5 ms·
when are we going to stop pretending these benchmarks have any meaning? anybody who's used these models knows that their real-world software engineering perfor
by notatoad 14d ago
when are we going to stop pretending these benchmarks have any meaning?
anybody who's used these models knows that their real-world software engineering performance has no relation to the ranking on deepSWE.
- caconym_ 14d ago+1. I've used the recent Gemini Flash models and I've used Opus 5, and the latter makes the former look like a box of broken crayons. Unless Flash 3.8 and/or this Muse Spark model are a much bigger deal than people seem to think, I will eat my hat if either one can come close to Opus 5 in actual real life "long-horizon software engineering" tasks. (I'm not happy about the above being true, but it's the reality I seem to inhabit.)
- jdm2212 14d agoAnd Fable 5.x makes Opus 5 look pretty dim, despite benchmarks suggesting they're comparable. The benchmarks really are just kinda meaningless.
- albrewer 13d agoA series of hot takes: Benchmarks are useful but only on a log2 basis. One model performing at 50% and another at 75% is just as impressive as one model performing at 78% and another at 90%. Confoundingly, a benchmark becomes useless once a frontier model scores over ~95% on them.
- jdm2212 13d agoI think that's definitely the right way to understand benchmark saturation, but there's a separate problem where the benchmarks are just not representative of real workflows even when they don't seem to be saturated.
- zackify 14d agoI have been using glm 5.3 flash and it feels as good as opus 5. Put a lot of work into it this week (100m tokens). Now I'm curious to try this one. These smaller models are getting very good imo
- KptMarchewa 13d agoNeither flash or regular glm 5.3 are close in my experience. I still prefer Sol though.
- zackify 12d agoWhat type of setup do you use? I have very small 4-5k initial context and do one task then clear. I rarely go above 150k context for most things.
- weiran 14d agoYep these software benches are only good at testing how well they can one shot. For the kind of attended/assisted development most of us do with agents it’s hard to find a benchmark that reflects my own experience of the frontier models still being quite far ahead.