Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
pich
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
20 ms
·
1.
▲
RSI AI Explained: What It Is and Where It Could Lead
(piszczek.pl)
3 points
by
pich
4d ago
|
0 comments
2.
▲
by
pich
1mo ago
A 3090 has 936 GB/s of bandwidth vs 432 GB/s on the 70W RTX PRO 4000 SFF, so 70 t/s there is not surprising. The interesting constraint here was fitting a workload-tuned 5.01 BPW quant + 256K + MTP into 24 GB while working wi
3.
▲
by
pich
1mo ago
Its not quite that simple here. The iMatrix-guided hybrid uses different quantization levels per tensor/layer, so there isnt one honest Q4/Q5/NVFP4 label I can put in the title
4.
▲
by
pich
1mo ago
5.01 BPW custom hybrid: bulk NVFP4, selected Q5_K/Q6_K tensors from an iMatrix, Q6_K embeddings and Q8_0 lm_head. The iMatrix was built from 5,472 messages across 296 real Hermes sessions
5.
▲
by
pich
1mo ago
vLLM is probably the key difference there… its scheduler is built around batching/concurrency, while this setup is heavily optimized llama.cpp for single-stream latency
6.
▲
Qwen3.8 27B at 256K: 50 TPS on a 24 GB GPU
(piszczek.pl)
38 points
by
pich
1mo ago
|
36 comments
7.
▲
DFlash changes what tokens per second means
(piszczek.pl)
2 points
by
pich
1mo ago
|
0 comments