Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
oshrimpton
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
oshrimpton
3mo ago
Yeah they are 100% in the wrong for removing the fine tuned codex models. It makes sense why they wouldn't want to allocate so many resources towards fine tuning but still the enshittification of GPT models is real
2.
▲
by
oshrimpton
3mo ago
I would be so curious to find a comprehensive benchmark on this, humans do have an unfortunate ahem Dunning-Kruger effect ahem tendency to do this
3.
▲
by
oshrimpton
3mo ago
Yeah the benchmark for sure isn't perfect and without super rigid prompting it is far too easy for it to get off course. 28% hallucination rate isn't nothing either
4.
▲
by
oshrimpton
3mo ago
Surprisingly not! It is the biggest hallucinator on the AA Omniscience Index just 2pp away from V4 Pro. I think this is partially due to the fact that Flash was trained on >32T tokens just like Pro deapite being almost 10x smaller - it s
5.
▲
by
oshrimpton
3mo ago
I'd definitely agree that it isn't directly model size, but there is the fact that a larger model in terms of parameter count needs a large amount of training data to not overfit or underfit. So I think this race to the top of &qu
6.
▲
by
oshrimpton
3mo ago
Agreed on the title, my bad! But yeah, I've had some truly terrible experiences using these "frontier" models in coding agents especially, where they just fabricate facts about codebases.
7.
▲
GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2
(arrowtsx.dev)
585 points
by
oshrimpton
3mo ago
|
294 comments