4 ms·
I quite wish they'd move to terminal bench 4.0. 2.1 is saturated - there is no world in which Gemini 3.8 Flash is producing better code than Astra or Fable, as
by dgacmu 12d ago
I quite wish they'd move to terminal bench 4.0. 2.1 is saturated - there is no world in which Gemini 3.8 Flash is producing better code than Astra or Fable, as the 2.1 results might suggest. The 4.0 results differentiate these models much more effectively.
(2.1 is useful for knowing they can all one-shot straightforward scripts, of course.)
- swingboy 11d agoDeepSWE has Gemini 3.8 Flash up really high, too.
- dgacmu 11d agoIt does and that one also feels kind of saturated for measuring the most advanced models - opus, Gemini, astra, sol, fable, glm, kimi all scoring within statistical noise of each other. (74 +-3% down to 69% +-5% for kimi). It's still providing strong discrimination between weaker models.