6 ms·
Time to dump this test. Probably not a coincidence every version has the same rolling green hills, gradient blue sky, sun in the corner, etc.
by pupppet 2mo ago
Time to dump this test. Probably not a coincidence every version has the same rolling green hills, gradient blue sky, sun in the corner, etc.
- fastball 2mo agoOn the one hand: yes, pelicans on bikes are definitely in the training set at this point. On the other hand: the test is clearly not saturated, given that you can see a clear difference in output at the various reasoning levels / model versions.
- pupppet 2mo agoI sort of agree, but within the same model I expect the reasoning effort to be reflected in the quality of output and that's basically how it played out. When you're comparing different models, then it's just who benchmaxxed the best and there's not a lot of value there.
- fastball 2mo agoBut that is my point: if benchmaxxing was all the labs were doing, then surely the dumber model could/would have equivalent performance? Rather than noticeably worse perf on a (somewhat trivial to game) test.
- dbt00 2mo agoGoodhart's law.
- modriano 2mo agoI don't know. If they were training on this, I feel like they would be able to get the shape of a bike frame right; it's a pretty simple polygon, and a lot of the bike frames that are getting generated would be impossible to steer.
- mbauman 2mo agoI mean, humans can't draw bikes! https://themagnet.substack.com/p/why-is-it-so-hard-to-draw-a-bike https://themagnet.substack.com/p/why-is-it-so-hard-to-draw-a...
- ceroxylon 2mo agoI partially agree, but in this case it kinda illustrates that it may not be worth using Terra on any reasoning level below high; those are some awful penguins on bikes.
- threatripper 2mo agoWaiting for the AMA on Reddit "Ten years ago I was responsible for the pelican department at OpenAI, AMA"