Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ehtbanton
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
LLM is a compiler, not a runtime
(getpocketbot.com)
3 points
by
ehtbanton
5mo ago
|
1 comments
2.
▲
by
ehtbanton
5mo ago
I just don't trust Anthropic's Claude Code team at all any more. Their tools are vibe-coded and their behaviour is anti-consumer. They shouldn't be surprised at the thousands moving to Codex every day.
3.
▲
by
ehtbanton
5mo ago
Benchmarks like this one are designed to thoroughly test the model across several iterations. 15% is a MASSIVE discrepancy. Come on Anthropic, admit what you're doing already and let us access your best models unhindered, even if it co
4.
▲
by
ehtbanton
5mo ago
This is genuinely very helpful. I'm planning a MacBook pro purchase with local inference in mind and now see I'll have to aim for a slightly higher memory option because the Gemma A4 26B MoE is not all that!
5.
▲
by
ehtbanton
5mo ago
This is very impressive, have tried it out. If only everyone was as good at making performant terminal applications (cough cough Anthropic)
6.
▲
by
ehtbanton
5mo ago
I've had this thought myself too. Going off on a slight tangent: I think there's also loads of useful stuff in domains like either of these which maps amazingly well to AI agent system design, but there's such a huge discrepa
7.
▲
by
ehtbanton
5mo ago
I will always maintain that the best benchmark is just trying it out for yourself. The most practical parallel for me is all the people posting about how some open-source model has "achieved X on Y benchmark - beating out Opus 4.6!&quo
8.
▲
by
ehtbanton
5mo ago
Wake me up when Anthropic does something right again...