7 ms·
I have no idea how artificial analysis got to be something anyone took seriously. This is their new benchmark set? A glance at their new index shows that whate
by throwaway13337 12d ago
I have no idea how artificial analysis got to be something anyone took seriously. This is their new benchmark set?
A glance at their new index shows that whatever they're measuring, it isn't useful.
Spend an hour with gemini 3.8 and tell me that model belongs in 2026. It feels like the model has Alzheimer's. It gets confused about whether what it reads is what it did. Just crazy bad.
I haven't tried muse spark 1.3. But it must have been a miracle since 1.2 to hit that rank.
Video game journalism vibes all over this.
- nojs 12d agoWhat other benchmarks do you recommend that are more accurate?
- WASDx 12d agoGive a task you have to 3 different models and see what actually works for you. There are no good benchmarks.
- dist-epoch 12d agoAs the saying goes, Artificial Analysis is the worst benchmarking company, except for all the others.
- Catloafdev 11d agoBecause it's the best option currently available. It's really easy to shit on AI benchmarks, but that noise is useless unless you're offering a solution or a better benchmark.