5 ms·
If indeed, as the new benchmarks suggest, this is the new "top dog" of models, why is the launch feeling a little flat? For comparison, the Claude 4 hacker new
by fumblebee 1y ago
If indeed, as the new benchmarks suggest, this is the new "top dog" of models, why is the launch feeling a little flat?
For comparison, the Claude 4 hacker news post received > 2k upvotes https://news.ycombinator.com/item?id=44063703 https://news.ycombinator.com/item?id=44063703
- Ocha 1y agoNobody believes Elon anymore.
- fumblebee 1y agoHm, impartial benchmarks are independent of Elon's claims?
- ben_w 1y agoImpartial benchmarks are great, unless (1) you have so many to choose from that you can game them (which is still true even if the benchmark makers themselves are absolutely beyond reproach), or (2) there's a difference between what you're testing and what you care about. Goodhart's Law means 2 is approximately always true. As it happens, we also have a lot of AI benchmarks to choose from. Unfortunately this means every model basically has a vibe score right now, as the real independent tests are rapidly saturated into the "ooh shiny" region of the graph. Even the people working on e.g. the ARC-AGI benchmark don't think their own test is the last word.
- irthomasthomas 1y agoIt's also possible they trained on test.
- bigyabai 1y ago"impartial" how? Do you have the training data, are you auditing to make sure they're not few-shotting the benchmarks?
- irthomasthomas 1y agoLikely they trained on test. Grok 3 had similarly remarkable benchmark scores but fell flat in real use.
- DonHopkins 1y agoThe latest independent benchmark results consistently output "HEIL HITLER!"
- brightfuturex 1y ago[dead]
- mppm 1y ago[flagged]
- Aerbil313 1y agoProbably more like Claude was slightly better than GPT-xx when the IDE integrations first got widely adopted (and this was also the time where there was another scandal about Altman/OpenAI on the front page of HN every other week) so most programmers preferred Claude, then it got into a virtuous cycle where Claude got the most coding-related user queries and became the better coding model among SOTA models, which resulted in the current situation today.
- v5v3 1y agoOther AI companies post a 5 minute article to read. This is a 50 minute long video, many won't bother to watch
- ceejayoz 1y agoI'm not sure there's any benchmark score that'd make me use a model that suddenly starts talking about racist conspiracy theories unprompted. Doubly so for anything intended for production use.
- typon 1y agoIts a shame this model is performing so well because I can't in good conscience pay money to Elon Musk. Will just have to wait for the other labs to do their thing.
- brightfuturex 1y agoI think it's a shame that your emotions are so much in your way. It's an illusion to think you can assess Elon at his true worth, like AI hallucinating due to lack of context.
- DonHopkins 1y ago[dead]
- fdsjgfklsfd 1y agoYou misspelled "principles".
- Kapura 1y ago[flagged]
- bilsbie 1y ago[flagged]
- johnfn 1y agoUpvotes are a lagging indicator. Despite all the leaderboard scores presented, etc, no one actually knows how good a model is until they go use it for a while. When Claude 4 got ~2k upvotes, it was because everyone realized that Claude 3.7 was such a good model in practice - it had little to do with the actual performance of 4.
- aprilthird2021 1y agoBecause the benchmarks are likely gamed. Also Grok had an extremely negative news cycle right before this, so the average bloke is skeptical that the smartest AI in the world thinks the last name Steinberg means someone is a shadowy, evil, cabal-type figure. Even though they aren't totally related, most people aren't deep enough in the weeds to know this