8 ms·
Low, High and Max, obviously, can't be compared across models. They only mean the model is likely to spend less reasoning effort (~output tokens) with Low than
by Palmik 1mo ago
Low, High and Max, obviously, can't be compared across models. They only mean the model is likely to spend less reasoning effort (~output tokens) with Low than High on the same, *single shot* task.
But even in this very post, you can see that Max was actually cheaper than High.
If you are using API, you should be comparing based on end-to-end cost or speed or whatever blend of those two matches your cost/time budget.