7 ms·
3.8 uses nearly twice as many tokens as 3.7. One might be inclined to think that they just increased the thinking budgets... 3.7 used 64M on high: https://arti
by eis 14d ago
3.8 uses nearly twice as many tokens as 3.7. One might be inclined to think that they just increased the thinking budgets...
3.7 used 64M on high: https://artificialanalysis.ai/models/gemini-3-7-flash https://artificialanalysis.ai/models/gemini-3-7-flash
3.8 used 120M on high: https://artificialanalysis.ai/models/gemini-3-8-flash https://artificialanalysis.ai/models/gemini-3-8-flash
Even their own chart showed more than 2x higher cost compared to 3.7: https://storage.googleapis.com/gweb-uniblog-publish-prod/images/gemini-3-8-cyber__evals__cwe-ben.width-2000.format-webp.webp https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...
- WASDx 14d ago3.7 high and 3.8 medium are essentially the same on AA intelligence and cost. Output tokens on DeepSWE gives the same picture. So there might be something to it but they have done other things as well. At least the tokens are really fast.
- zuzululu 14d agoi find deepswe not very reliable for instance it puts grok 4.6 xhigh over sol medium