Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
somnial
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
What happens when a GPU writes memory
(blog.doubleword.ai)
5 points
by
somnial
17d ago
|
0 comments
2.
▲
What happens when a GPU reads memory?
(blog.doubleword.ai)
15 points
by
somnial
1mo ago
|
0 comments
3.
▲
The case for disaggregated LLM serving
(blog.doubleword.ai)
4 points
by
somnial
1mo ago
|
0 comments
4.
▲
On-the-fly snapshot compression for elastic inference at scale
(blog.doubleword.ai)
6 points
by
somnial
1mo ago
|
1 comments
5.
▲
by
somnial
1mo ago
throughput scales superlinearly with number of GPUs when networked well and deployed with wideEP, so 1x won't compare. also it would be interesting to figure from the DSpark paper whether their numbers are consistent with the GPUs stil
6.
▲
NVLink, NVSwitch, and All That
(blog.doubleword.ai)
5 points
by
somnial
2mo ago
|
2 comments
7.
▲
The Anatomy of an Instruction Pipeline Hazard
(hiraditya.github.io)
9 points
by
somnial
2mo ago
|
0 comments
8.
▲
Width vs. Depth: Speculating on the Margin
(blog.doubleword.ai)
17 points
by
somnial
3mo ago
|
1 comments
9.
▲
by
somnial
3mo ago
this is a blog post from a company that hosts open weights LLMs ( https://www.doubleword.ai/ ). I think its possible it might have been tongue in cheek
10.
▲
by
somnial
3mo ago
https://fergusfinn.com/blog/economics-of-speculative-decodin... good point tho - plus for Deepseek the shared expert increases the overlap slightly
11.
▲
by
somnial
3mo ago
true, but no reason the predictor model couldn't use linear attention (i.e. mamba, GDN etc) to predict KV caches
12.
▲
Pushing memory bound CUDA kernels past the speed of light with data compression
(fergusfinn.com)
2 points
by
somnial
4mo ago
|
0 comments
13.
▲
Speculative KV coding: ~4× losslessly compressed KV cache using a small model
(fergusfinn.com)
2 points
by
somnial
4mo ago
|
0 comments
14.
▲
70x faster cold(ish) starts for SGLang
(fergusfinn.com)
1 points
by
somnial
5mo ago
|
0 comments
15.
▲
LLM powered data structures: A lock-free binary search tree
(fergusfinn.com)
1 points
by
somnial
8mo ago
|
0 comments
16.
▲
Parallel Primitives for Multi-Agent Workflows
(fergusfinn.com)
1 points
by
somnial
9mo ago
|
0 comments
17.
▲
Scheduling in LLM Inference
(fergusfinn.com)
1 points
by
somnial
10mo ago
|
0 comments
18.
▲
How fast can an LLM go?
(fergusfinn.com)
2 points
by
somnial
11mo ago
|
0 comments