7 ms·
If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3
by postalcoder 6d ago
If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%).
Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
- enraged_camel 6d agoYeah, this echoes my thoughts. I will be very surprised if a model with 2.8T parameters reaches the intelligence and capabilities of 10T parameter models. RL can take things far, but not that far.
- nullbio 6d agoClosed weights AND benchmaxxed. Somehow this company raised 2bil at a 48bil valuation. Pure insanity. I feel bad for their investors (not really, but... Still). Andreessen Horowitz is being played like a fiddle.
- thereitgoes456 6d agoThe Cursor acquisition shows that it’s possible for these valuations to be justified. But Cursor was more successful and bent the truth much less. While I wouldn’t expect anything good for Cognition’s fate, it’s a much safer bet than Thinking Machines, SSI, and some others. Though they’ll be in big trouble if the more talented Chinese labs stop letting them repackage their work.
- selectodude 6d agoThe cursor acquisition just shows that there’s always a dumber shithead out there. Though vaporizing Elon Musks money is about as pure of a good as there is out there these days.
- throwaway240403 6d agoImprovements in models and products coming out of SpaceXAI since the acquisition would seem to disagree with you.
- throwup238 6d ago> Andreessen Horowitz is being played like a fiddle. Andreessen Horowitz is not being played like a fiddle here. This might be their only investment in a decade that isn’t entirely predicated on being a scam.
- deleted 6d ago[deleted]
- mediaman 6d agoThis is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.
- ActionHank 5d agoTB4 has not been saturated yet. They are all gaming these benchmarks, it is perfectly reasonable not to trust any of them.
- felixgallo 6d agoIs Sol benchmaxxed? Of course it is. Altman was caught in previous attempts trying to game benchmarks, does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with?
- general_reveal 6d ago[flagged]
- moomin 6d agoA lot of people seem to be the Antichrist these days. I’d have though we’d get a lull after the millennium but it’s all the rage.
- ZeWaka 6d agoI guess we also need more antipopes.
- general_reveal 6d agoWell, the antichrist should be in and around this AI thing for one particular reason: The devil cannot create anything of his own because he is not God, by definition. We have already observationally defined generative AI as something that cannot create anything novel in the sense it cannot output anything it has never seen (cannot create new, always a re-assortment of what is). In that way , AI is a perfect mimicry of how the devil operates (in totality, as the devil perverts and replicates anything good, often subtly and always deceptively), which is to thieve off God, steal. So he would be around, if you catch my drift, right about now. And I wouldn’t be shocked if he’s on HN, and that he would chose technology as the vessel. And ultimately, when it’s all said and done, I would not be shocked that those who studied and developed AI, did so for the devil whether they were aware or not. Anyway, let a poor Christian have his end-times hypothesis.
- eranation 6d agoWhen a benchmark becomes a target, it's no longer a good benchmark...
- tonychang430 6d agopeople are just fighting for numbers.. i don't fundamentally see the model being better
- fallingbananna 6d agoThose 27.3% are still in the ballpark of modern models: - Sonnet 5 - 12.4% - Luna - 17.3% - Grok 4.6 - 20.3% - Sol - 37.3% - GLM 5.3 - 41.8% - Opus 5 - 51.8%
- Readerium 6d agoDeepSeek v4.1 Flash 31.2% Source: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash#comparison-with-frontier-models-max-reasoning-effort https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash#compa...
- p1esk 6d agoAstra is 58%. The current title says it's "rivaling Astra"
- thereitgoes456 6d agoIt is rivaling Astra, on their own benchmark that they made (FrontierCode), that they ran themselves in their own closed-source ecosystem that isn’t reproducible by anyone.
- nijave 6d agoThis explains a lot about Sonnet 5.
- mokre 6d agoGLM 5.3 looks strange, because of this Chinese labs benchmaxx moto. So rather they have emergent abilities or... Also a lot of questions to benchmark because opus 5 is completely useless model right now. I think that the main problem with opus that they try to solve context size optimization problem, and that is the main reason why it speaks like alien with only one technical dictionary at hand. So why it is so good?
- throwatdem12311 6d agoThis is why I find benchmarks absolutely worthless. First, almost all models are within spitting distances of eachother. Second, it never translates to being better for my own workloads. You just need to make your own benchmarks.
- walrus01 6d agoFor comparison Qwen 3.8-Flash-Next which runs in under 190GB of RAM locally scores 25.3% on terminalbench 4.0.
- Readerium 6d agoYeah DeepSeek V4.1 beats this by 15 percent (4 points) on terminal bench 4.0
- thefourthchime 6d agoCame here to say the exact same thing! People have to stop paying any attention to coding benchmarks that aren't Terminal Bench 4. I noticed I noticed they didn't include Gemini 3.8, which also murders DeepSWE and Terminal Bench 2.0 -- because they are useless benchmarks now! Of course in a couple months TB4 will also be old hat, so TB5 will have to be the new real benchmark.
- dudeinhawaii 6d agoYour post made me wonder if Artificial Analysis had finally moved to TB4 and lo and behold they have and Astra is tied with Fable 5.1 at 53. That then made me realize that they lower the bars of tied scores so on the site it looks like Astra in second place. Weird. Anyway, yes, so many of these composite benchmark sites are irrelevant if they're not trimming the fat and sticking to the most up-to-date variants.
- muddi900 5d ago> benchmaxxed An aside: When did talking like incels became cool?