Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kkm
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
Claude.md is good for taste and project context.It's a weak place for invariants
(tesseracted-labs-blog.vercel.app)
3 points
by
kkm
2d ago
|
0 comments
2.
▲
Moving coding-agent guardrails from prompts to hooks
(tesseracted-labs-blog.vercel.app)
4 points
by
kkm
3d ago
|
0 comments
3.
▲
Improving Throughput by Optimising KV Cache Efficiency for Agentic Workloads
(j9s.io)
3 points
by
kkm
4d ago
|
1 comments
4.
▲
Guide to the Kimi DeltaNet Family of linear attention
(blog.doubleword.ai)
3 points
by
kkm
2mo ago
|
0 comments
5.
▲
Forensic Analysis of Container Snapshot Chains for Post-Event Reconstruction [pdf]
(radostin.io)
2 points
by
kkm
2mo ago
|
0 comments
6.
▲
The Agent swarm that designs itself
(peterbhabra.com)
1 points
by
kkm
3mo ago
|
0 comments
7.
▲
Don't Build a Router. Train the Small Model to Know When to Defer
(distillabs.ai)
2 points
by
kkm
3mo ago
|
1 comments
8.
▲
The gap between open weights LLMs and closed source LLMs
(blog.doubleword.ai)
306 points
by
kkm
3mo ago
|
250 comments
9.
▲
InfiniBand, RoCE, and All That
(fergusfinn.com)
5 points
by
kkm
3mo ago
|
0 comments
10.
▲
2678x Faster Matrix Multiplication with a GPU
(0mean1sigma.com)
2 points
by
kkm
3mo ago
|
0 comments
11.
▲
UCCL-EP: DeepEP-style expert parallelism on any NIC, no GPU-initiated comms
(fergusfinn.com)
9 points
by
kkm
3mo ago
|
0 comments
12.
▲
Hacking Google with A.I. For $500k
(brutecat.com)
1 points
by
kkm
3mo ago
|
0 comments
13.
▲
How to setup a local coding agent on macOS
(ikyle.me)
507 points
by
kkm
3mo ago
|
127 comments
14.
▲
Anatomy of a high-performance EP kernel
(fergusfinn.com)
16 points
by
kkm
3mo ago
|
1 comments
15.
▲
No Token Left Behind: Demystifying Token-in-Token-Out in Miles
(lmsys.org)
2 points
by
kkm
3mo ago
|
0 comments
16.
▲
MoE expert co-activations: Reordering inputs yields easy throughput gains
(blog.doubleword.ai)
2 points
by
kkm
3mo ago
|
0 comments
17.
▲
The Economics of Speculative Decoding
(fergusfinn.com)
30 points
by
kkm
3mo ago
|
6 comments
18.
▲
Speculative KV coding: losslessly compressing KV cache by up to ~4×
(fergusfinn.com)
155 points
by
kkm
4mo ago
|
48 comments
19.
▲
70x faster cold(ish) starts for SGLang
(fergusfinn.com)
1 points
by
kkm
4mo ago
|
0 comments
20.
▲
by
kkm
4mo ago
This is very interesting, planning to write about it?
21.
▲
by
kkm
4mo ago
Also the vllm patch accompanying the blogpost: https://github.com/doublewordai/vllm-amd-blog-doubleword
22.
▲
Bringing Up DeepSeek-V4-Flash on AMD MI300X
(fergusfinn.com)
120 points
by
kkm
4mo ago
|
25 comments
23.
▲
Brave AI privacy:LLMs on NEAR AI Nvidia-Backed Trusted Execution Environments
(brave.com)
1 points
by
kkm
10mo ago
|
0 comments
24.
▲
How fast can an LLM go?
(fergusfinn.com)
2 points
by
kkm
10mo ago
|
0 comments
25.
▲
FHE can be leveraged for LLMs such as ChatGPT in a privacy-preserving manner
(huggingface.co)
4 points
by
kkm
2y ago
|
0 comments
26.
▲
Harnessing the Power of Large Language Models for Insightful Review Analysis
(techblog.holidaycheck.com)
1 points
by
kkm
2y ago
|
0 comments
27.
▲
A Privacy-First approach to use AI for understanding our Customers Better
(techblog.holidaycheck.com)
1 points
by
kkm
3y ago
|
0 comments
28.
▲
How to make LLMs go fast
(vgel.me)
2 points
by
kkm
3y ago
|
0 comments
29.
▲
Leveraging Large Language Models for Sentiment Classification in Hotel Reviews
(techblog.holidaycheck.com)
1 points
by
kkm
3y ago
|
0 comments
30.
▲
Managers Should Think More Like Hackers
(hbr.org)
3 points
by
kkm
3y ago
|
0 comments
More ›