30 ms·
I initially had unbelievably terrible experiences with Opus 5 and Fable in their higher reasoning levels. I've had WAY better results on medium effort. IIUC,
by onlyrealcuzzo 25d ago
I initially had unbelievably terrible experiences with Opus 5 and Fable in their higher reasoning levels.
I've had WAY better results on medium effort.
IIUC, the consensus seems to be that anything more than medium effort is rarely worth it - and you far more often run into these extreme worst cases than you do with even the lowest effort levels. That definitely coincides with my anecdata.
It's really only worth it if you're hoping to win the lottery asking it to solve an Erdos problem.
- gwerbin 25d agoI found that I need the big model and the high reasoning effort on tasks where I have a large amount of details of varying levels of importance to keep in mind, all affecting different aspects of the project that might be interrelated to various degrees. Anything less and it would lose track of details. Whereas the very large amount of thinking tokens seem to give the model a chance to "remember" everything it needed to in order to produce good output. For example I'm working on a project now where I need to keep in mind details from 3 separate source repos in different programming languages, along with probably a dozen important business-related documents that either corroborate the stuff in the source repos or add additional important context. It's a lot of details for even a human to manage, and when it comes to actually synthesizing plans and reports across this sprawling information environment, anything less than Opus 5 on Extra+ tends to miss important details and make bad recommendations or draw incorrect conclusions, which then poison subsequent context. I suspect some kind of RAG-like memory system would greatly facilitate a project like this, but even with such a system I'm not confident that I could get away with an LLM that "thinks" less hard than this. I will say that slogging through the generated documents is kind of miserable and I have to repeatedly fork off side conversations to ask for clarification, but sometimes leads the model to "realize" it's made a mistake in all of its dense babbling, and it's all very hard to interpret. I never had much interest in trying GPT 5.6 until now.