4 ms·
Annoying they didn't show benchmarks for several effort modes, since it seems like it might close the gap with Opus 4.8 by cranking tokens up? Noam Brown (Open
by conradkay 2mo ago
Annoying they didn't show benchmarks for several effort modes, since it seems like it might close the gap with Opus 4.8 by cranking tokens up?
Noam Brown (OpenAI) "Implications of Large-Scale Test-Time Compute" https://xcancel.com/i/article/2064210146558136827 https://xcancel.com/i/article/2064210146558136827