5 ms·
Yes, they are all benchmaxxed, but the question is how benchmaxxed they are relative to each other. We run an evaluation that only compares models in open-ende
by gertlabs 2mo ago
Yes, they are all benchmaxxed, but the question is how benchmaxxed they are relative to each other.
We run an evaluation that only compares models in open-ended multi-agent environments where agents affect each other, primarily testing writing code. It's designed to be less vulnerable because there's no solution set, and it's been pretty effective and tends to rank Chinese models lower than their advertised model cards (relative to US models). Kimi K3 is a bit of an exception there -- it truly is a near-frontier model. But it's so slow. Ironically, Muse Spark 1.1 is one of the strongest models we've tested after Fable and Sol while also leading the cost efficiency curve. Big turnaround from Llama 4.
Data at https://gertlabs.com/rankings https://gertlabs.com/rankings
- isityettime 2mo agoAlthough you mention cost, I don't see task cost or total evaluation cost per model in your data. Am I just missing it?
- gertlabs 2mo agoThe Efficiency tab at https://gertlabs.com/rankings?mode=oneshot_coding https://gertlabs.com/rankings?mode=oneshot_coding (only have cost data for the coding evaluations)
- isityettime 2mo agoThanks!