5 ms·
You are very much not alone, and I don't think we're all getting dumber -- I kept using opus 4.5 all through the nonsense that was 4.7, 4.8, and 5, and kept hav
by pickledish 27d ago
You are very much not alone, and I don't think we're all getting dumber -- I kept using opus 4.5 all through the nonsense that was 4.7, 4.8, and 5, and kept having a good time :)
I think we just need to decouple "doing better on benchmarks" and "actually more useful to me", since they've clearly diverged
- serf 27d agodoesn't that just mean that either a) you're using the wrong benchmark to judge or b) the benchmark that YOU need doesn't exist.
- gwerbin 27d agoThat's not the point, the point is that the company making the product is optimizing for the benchmark and/or the apparently idiosyncratic preferences of their own team, and not for the user experience of their paying customers.
- rkuodys 27d agoCompany can optimise for the benchmark (profit) while worsening the product. I think thebterm enshitification is used there. It appears that AI got it too
- ethbr1 26d ago> I think we just need to decouple "doing better on benchmarks" and "actually more useful to me", since they've clearly diverged There's a third perspective here: models are getting less useful, but overfitting to seeming useful to humans. Imho, this is why analysis like TFA + third party cross-compatible harnesses (read: last mile UX) are so important to the leading labs optimizing for actual utility. I'm suspicious enough of my subjective evaluation to believe a well-designed harness / verbiage could gaslight me into believing an objectively inferior model was superior. And at some point frontier labs are looking at the ROI of investing $1 in that vs actual model improvement.