8 ms·
I see no point having these pelicans used for anything related model qualification.
by m00dy 14d ago
I see no point having these pelicans used for anything related model qualification.
- simonw 14d agoAt this point the only thing they're useful for is visualizing the differences between effort levels and roughly tracking the progression of models within a specific model family. And they still do that really well!
- menaerus 14d agoI don't see how useful this benchmark at all is for tracking the progression of models. I am not intending to bash on you personally but this is useless. People who are using AI models everyday are for sure not interested how close the AI model can visualize the pelican but they are interested in how they will perform on their daily tasks at work or private use. Correlation between doing good on pelican task and doing good on actual work you need to do is close to zero.
- simonw 14d agoLook at the difference between the Muse 1.2 and Muse 1.3 results.