Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
brausepulver
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
brausepulver
8d ago
You need to separate memory capacity and bandwidth. Looping decreases memory capacity/FLOP but not bytes loaded/FLOP, since weights need to be loaded again for the 2nd pass. Plus (depending on the method used) capacity required fo
2.
▲
by
brausepulver
29d ago
What LLM-specific hardware improvements should one expect? Seems to me that LLM inference is simple architecturally (matmul et al) so most scaling in hardware should come from general improvements (memory BW, packaging, interconnect, power)
3.
▲
by
brausepulver
2mo ago
Consider that a lot of the resistance to just existing is external expectations. I feel like I should always be doing something because I'm pressured to, or that small things bother me because I feel I need to perform. Being content