Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
anana_
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
anana_
5d ago
I'm curious what the actual rate of replacement/useful life of these GPUs running AI inference 24/7 is. If these cards burn out in less than the ~5 years of depreciation that accounting puts them at, well then there will be p
2.
▲
by
anana_
1mo ago
Agreed. Luckily, this model also scores high in AA non-hallucination, so it knows what it doesn't know -- perfect for situations where it can just tool call a web search.
3.
▲
by
anana_
1mo ago
And to read the tea leaves a little: 3.8 actually performs slightly worse than 3.6 on AA-Omniscience Accuracy, which could imply that they traded out world knowledge for capability in other areas. It also produces nearly twice as many tok
4.
▲
by
anana_
1mo ago
For more context, this puts it on par with models like GLM 5.2 and GPT 5.6 Luna, which are far larger
5.
▲
Qwen3.8 27B scores 52 on Artificial Analysis
(artificialanalysis.ai)
381 points
by
anana_
1mo ago
|
180 comments
6.
▲
by
anana_
1mo ago
When stuff like this: https://doublespeed.ai/ exists I don't find that hard to believe at all, although it cuts both ways
7.
▲
by
anana_
1mo ago
Seems like MTP is available immediately!
8.
▲
by
anana_
1mo ago
As was the case with GLM 5.3, it seems that there is still much juice to be squeezed from post-training
9.
▲
by
anana_
1mo ago
Monstrous benchmarks! Hoping it is not benchmaxxed.
10.
▲
by
anana_
1mo ago
What a week for AI model releases
11.
▲
by
anana_
2mo ago
Apparently agentic performance in Gemma was improved recently: https://x.com/googlegemma/status/2077449152062247219 Too little too late imo
12.
▲
by
anana_
3mo ago
They keep mentioning a 31B dense model, but there are no benchmarks or weights for it anywhere?
13.
▲
by
anana_
3mo ago
It looks like the purpose of this model is to i. generate environmental sim data for doing RL on other models or ii. act as a foundation model (they trained it to select actions as well as predicting the next state in the same loop?) Either
14.
▲
by
anana_
3mo ago
I believe the benchmark listed is about simulating the environment for the various tasks, rather than doing them. It seems that the point of this model is to generate sim data to improve other models with
15.
▲
by
anana_
3mo ago
Unfortunately on Strix Halo or any similar unified memory set up, dense models are gonna be dirt slow due to the tiny memory bandwidth... But I agree, 27B is superior.
16.
▲
by
anana_
3mo ago
Perhaps try a different model? Just from anecdotal experience, I find that the Gemma models smaller than 31B do not tool call as often as they should. Some of the benchmarks appear to back this up [0] Of course, a lot depends how you are us
17.
▲
by
anana_
3mo ago
I have one too and it never occurred to me to use it for anything other than games. Would be interested in seeing how you did it!
18.
▲
by
anana_
5mo ago
They do now - https://support.mozilla.org/en-US/kb/use-sidebar-access-tool...
19.
▲
by
anana_
5mo ago
Hypothesizing here, but maybe the idea is sort of a form of technological/economic warfare? Releasing performance equivalent yet more cost efficient open weight models should in theory drive the cost of inference down everywhere. This
20.
▲
by
anana_
5mo ago
I'm not saying it's the latest Qwen iteration - that would be Qwen3.6. I'm saying it's the latest iteration of the finetuned model mentioned in the parent comment. I'm also not suggesting that it's "the la
21.
▲
by
anana_
5mo ago
It's rather surprising that a solo dev can squeeze more performance out of a model with rather humble resources vs a frontier lab. I'm skeptical of claims that such a fine-tuned model is "better" -- maybe on certain benc
22.
▲
by
anana_
5mo ago
Upon rereading, I'd agree. Fits with the tone of the rest of the write up.
23.
▲
by
anana_
5mo ago
> Sometimes you need the absolute cutting-edge reasoning of Claude 3.5 Sonnet or GPT-4o Dead giveaway
24.
▲
by
anana_
7mo ago
https://huggingface.co/Qwen/Qwen3.5-27B I wasn't aware of that, which page mentions that?
25.
▲
by
anana_
7mo ago
I've had even better results using the dense 27B model -- less looping and churning on problems
26.
▲
by
anana_
7mo ago
They own GEICO...