Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kroaton
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
kroaton
4d ago
You need to validate what providers are actually serving. Add benchmarks, properly showcase what quantization they are serving on the model and KV cache, etc. Until that happens, your service is doomed to be shitty.
2.
▲
by
kroaton
7d ago
I've had Astra say that it is a Qwen model. They are all cross-trained and distill each other.
3.
▲
by
kroaton
7d ago
I wouldn't trust OpenCode Go with my lunch after the shit-sandwich they served us with horrible V4 quants and 0 transparency. Then blaming it on their partners.
4.
▲
by
kroaton
7d ago
Considering the fact that Google/Anthropic/OpenAI have WAY more compute and the race is this close, it's obvious that DeepSeek/GLM/Qwen teams are better or we're approaching a wall in terms of progress.
5.
▲
by
kroaton
10d ago
Garbage.
6.
▲
by
kroaton
13d ago
I think it mostly shows that there is no moat and the only advantage the U.S companies have over the Chinese is more compute. Qwen Max, Kimi K3, GLM 5.3 are really close to Opus/Sol/Fable/Astra and they are open weights.
7.
▲
by
kroaton
13d ago
The only good take here.
8.
▲
by
kroaton
13d ago
Especially since they still serve Codex-Spark, which is dogshit.
9.
▲
by
kroaton
22d ago
But Bonsai is garbage.
10.
▲
by
kroaton
23d ago
As if our EU leaders aren't a complete joke as well. Pushing ChatControl, fascism and gambling everywhere. We're just as much of a joke.
11.
▲
by
kroaton
1mo ago
This will degrade performance significantly. LLama.cpp has had this for a while and it tanks benchmark performance. I ran GPQA on GLM 5.2 using the llama implementation and it came back 19 points under the regular results.
12.
▲
by
kroaton
1mo ago
Or you can use parlor to chat with it directly https://github.com/fikrikarim/parlor/
13.
▲
by
kroaton
2mo ago
Better than Opus 4.8 on complex tasks but tends to overthink. It found a bunch of bugs and architecture issues that only 5.6 Sol Max and Fable on my C++ projects.
14.
▲
by
kroaton
2mo ago
It depends on what you need, but Krea/Klein9b/Ideogram4/Z-Image are among the best right now for text2image and Qwen Edit and Klein are probably still the best at editing.
15.
▲
by
kroaton
2mo ago
Yup. Smells like marketing.
16.
▲
by
kroaton
2mo ago
They've been investing heavily into this over the last 2 releases, but it's just that SideFX is a really small company. I think they have 50 people or so right now. Karma is way more accurate than Cycles, but slow. Karma XPU is a
17.
▲
by
kroaton
2mo ago
And that's when you learn Houdini and open up a workflow from 2008 and it still works the same, despite SideFX pushing out an insane amount of features every release (look at their SneakPeaks).
18.
▲
by
kroaton
2mo ago
Work more for my boss and landlord so the fiefdom can survive.
19.
▲
by
kroaton
2mo ago
It also goes to show that Fable/Sol must be 4-5T in size.
20.
▲
by
kroaton
2mo ago
Ling/Ring 1T-A50B and the new Inkling 975B-A41B deserve to be on that list.
21.
▲
by
kroaton
2mo ago
Which would be very interesting to test, as larger models (such as Deepseek V4 Flash or Qwen 397B) seem to compress better. Their Q2 quants are usable as is, even without the ternary compression.
22.
▲
by
kroaton
2mo ago
A shame it has Marc Maron in it.
23.
▲
by
kroaton
3mo ago
Please don't use that garbage. Just use the base Qwen models or Nex/Orinth, as those are the only properly post-trained finetunes. The Qwopus models are marketing.
24.
▲
by
kroaton
3mo ago
"DeepSeek-V4-Flash will fit" At Q2, 2bit? Lobotomized to death.
25.
▲
by
kroaton
3mo ago
NVFP4 will be better if the model provider actually post-trained properly after quantizing.
26.
▲
by
kroaton
3mo ago
You're late to the party, mate; we've been doing this for years. Grab a SearXNG instance, stand up an MCP server for it, and expose the tool into your system prompt. Or use Brave Search. Or Exa if you want to pay. Any of them work
27.
▲
by
kroaton
3mo ago
https://github.com/nexu-io/open-design
28.
▲
by
kroaton
3mo ago
This has been answered many times already. The short version: Brave's adblocker is not an extension, so manifest v3 has no effect on it.
29.
▲
by
kroaton
3mo ago
To be fair, GPT5.5-Xhigh is similarly capable and has not burned the world down.
30.
▲
by
kroaton
4mo ago
Anthropic nuked a big chunk of that "developer sentiment" when they rug-pulled us with the rate limits and gaslit us with "it was just a bug, guys!".
More ›