6 ms·
In theory that's what benchmarks are for. If you're assuming they're "benchmaxxed", note that new benchmarks have been released after the model came out that it
by usef- 1mo ago
In theory that's what benchmarks are for. If you're assuming they're "benchmaxxed", note that new benchmarks have been released after the model came out that it did well on without being trained.
Do you have any links to credible claims or independent benchmarks that found they were a step down? Or a specific task that worked worse for you?
My private benchmark tasks, and independent evaluators I've seen all overwhelmingly showed improvement.
Every model released for the past four years has had claims on the internet of getting worse. But transcripts are permanent so it should be easy to give a side by side of an earlier task that is now worse. I don't ever see people do that. Instead I see that every single task on a computer that is verifiable is now night-and-day better.
I'm genuinely curious if you've used them yourself or you're judging this based on internet commentary?