Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
corsix
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
corsix
1y ago
In this case, "atomic 128-bit store" is the special instruction, with the twist that half of those 128 bits contain a pointer.
2.
▲
by
corsix
1y ago
You can start with this idea, and then make it _very_ performant by using bit counting instructions. See https://www.corsix.org/content/higher-quality-random-floats for an exposition.
3.
▲
by
corsix
2y ago
Indeed, LuaJIT support for Windows/Arm64 was added in https://github.com/LuaJIT/LuaJIT/issues/593 , and there’s experimental out-of-tree support for Windows/Arm64EC in https://github.com&#
4.
▲
by
corsix
2y ago
For an implementation of logical immediate encoding without the loop, see https://github.com/LuaJIT/LuaJIT/blob/04dca7911ea255f37be799...
5.
▲
(Ab)using gf2p8affineqb to turn indices into bits
(corsix.org)
2 points
by
corsix
2y ago
|
0 comments
6.
▲
Polyfill-Glibc
(github.com)
5 points
by
corsix
2y ago
|
1 comments
7.
▲
by
corsix
2y ago
AArch64 NEON has the URSQRTE instruction, which gets closer to the OP's question than you might think; view a 32-bit value as a fixed-precision integer with 32 fractional bits (so the representable range is evenly spaced 0 through 1-ε,
8.
▲
by
corsix
3y ago
Unfortunately things aren't so simple, as when doing JIT compilation, LuaJIT _will_ try to shorten the lifetimes of local variables. Using the latest available version of LuaJIT ( https://github.com/LuaJIT/LuaJIT&#x
9.
▲
by
corsix
3y ago
Some compute might be on the AMX units (dedicated matrix multiplication coprocessor, closely attached to the CPU, distinct from both ANE and GPU). They gained bf16 support in M2.
10.
▲
by
corsix
3y ago
close() is documented as a cancellation point, and is the kind of syscall that might crop up in a destructor.
11.
▲
by
corsix
3y ago
Using your cdf framing, while there is a point about [0, 1) versus (0, 1] intervals, the bigger point of the article is about whether said cdf holds for any IEEE-754 double-precision p, or whether it only holds for p of the form i*2^-53 (fo
12.
▲
by
corsix
3y ago
Oh, cute. It looks like they ever-so-slightly overweight the probability of values whose mantissa is entirely zeroes though. For example, the probability of hitting exactly 0.5 should be 2^-54 + 2^-55, whereas zig looks to give it 2^-54 + 2
13.
▲
by
corsix
3y ago
https://www.corsix.org/content/whirlwind-tour-aarch64-vector... is my take on NEON, albeit not quite the same form factor as the OP.
14.
▲
by
corsix
3y ago
I think https://github.com/LuaJIT/LuaJIT/commit/6a2163a6b45d6d251599... improved things a bit, notably making automatic tarballs work again.
15.
▲
by
corsix
3y ago
I’m pleased to see FUTEX2_SIZE_U64, but saddened that it isn’t actually implemented. It has always seemed like a very useful primitive to have.
16.
▲
by
corsix
3y ago
Also interesting is that https://luajit.org/status.html now states “LuaJIT is actively developed and maintained” (whereas for the last ~5 years, “actively” isn’t a word I’d have used), and makes reference to a TBA developme
17.
▲
by
corsix
3y ago
Per https://www.stateof.ai/compute , one of the players in the market has ten thousand GPUs in a private cloud. Out-computing just that one player is hard enough, let alone out-computing the whole market.
18.
▲
by
corsix
3y ago
The trie structure described in the article can be (ab)used to export an infinite number of symbols from a library: https://www.corsix.org/content/exporting-an-infinite-number-...
19.
▲
by
corsix
3y ago
From a hardware perspective, vector instructions operate on small 1D vectors, whereas tensor instructions operate on small 2D matrices. I say “instructions”, but it’s really only matrix multiply or matrix multiply and accumulate - most othe
20.
▲
by
corsix
3y ago
Assuming that you're after "round to nearest with ties toward even", then the quoted numpy code gets very close to `vcvtps2ph`, and one minor tweak gets it to bitwise identical: replace `ret += (ret == 0x7c00u)` with `ret |=
21.
▲
by
corsix
3y ago
The first niche that came to mind was x86 code running under Rosetta 2; despite ARM having an equivalent to F16C, Rosetta 2 doesn’t translate AVX, and F16C doesn’t have a non-AVX encoding.
22.
▲
by
corsix
4y ago
Using the notation from the article, N+K is sufficient for RS(N,K). One point of confusion is that different authors use different notation; some use RS(num data shards, num parity shards), some use RS(total num shards, num data shards), an
23.
▲
by
corsix
4y ago
The next article in blog order is one application: https://www.corsix.org/content/reed-solomon-for-software-rai... Another application is crypto: the SubBytes step of AES maps very neatly onto gf2p8affineinvqb, so algo
24.
▲
by
corsix
4y ago
Alternatively, the following pair of articles, the first of which is already referenced as a footnote in the OP: http://www.corsix.org/content/galois-field-instructions-2021... http://www.corsix.org/con
25.
▲
by
corsix
4y ago
Lua gets this right - the lowering of loops (e.g. https://www.lua.org/manual/5.1/manual.html#2.4.5 ) says “var is invisible” and has “local v = var”, the latter akin to Go’s “item := item”
26.
▲
by
corsix
4y ago
Performance comparison is hard given that the Intel one hasn’t shipped yet.
27.
▲
by
corsix
4y ago
The elephant in the room contrast: Apple has been shipping this for years, whereas Intel _might_ ship theirs in something later this year.
28.
▲
by
corsix
4y ago
Slightly cursed is how I roll. I would like to know where the trick first originated from though (I found it at https://github.com/yvt/amx-rs/blob/main/src/nativeops.rs#L22 rather than inventing it
29.
▲
by
corsix
4y ago
https://www.corsix.org/content/contrasting-intel-amx-and-app... compares this with Intel’s AMX
30.
▲
Apple AMX Instruction Set
(github.com)
6 points
by
corsix
4y ago
|
0 comments
More ›