5 ms·
There is plenty of evidence that they have improved in all benchmarks and also in my private experience. But have they improved in the things they still fail at
by threatripper 14d ago
There is plenty of evidence that they have improved in all benchmarks and also in my private experience. But have they improved in the things they still fail at? No, they still fail at them. You need only one example of failure to prove that it still fails. They still fail a lot on many real world tasks.
So, depending on what you ask, they may have not improved even a tiny bit.
- Denkel 11d agoI am a very light user, so this is my feeling from reading about other people's experience; I wouldn't say that they plateau'd but up to 4.5/4.8 the gains in the models felt exponential while since then they feel more linear and the big improvements are coming less from the models and more from everything around it (harnesses, agentic development, skills...). So, while I don't feel like there has not been improvement, it really feels like there is a limit that will be reached sooner than later (and for sure before any AGI).