5 ms·
Currently top at https://deepswe.datacurve.ai https://deepswe.datacurve.ai - beating Opus 5! https://artificialanalysis.ai/models/gemini-3-8-flash https://arti
by mattlondon 16d ago
Currently top at https://deepswe.datacurve.ai https://deepswe.datacurve.ai - beating Opus 5!
https://artificialanalysis.ai/models/gemini-3-8-flash https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium!
Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.
- sunaookami 16d ago>shows an intelligence score of 59, the same as Opus 5! ...on Medium reasoning. Claude Opus 5 (high) is the default in e.g. Claude Code and scores 61. Still very impressive.
- Gecko4072 16d agoGoogle - we're so back
- oceanplexian 16d agoOnly 1 point behind the Chinese SOTA from two months ago.
- roosterIllusi0n 16d agoI had qwen 3.8 3bit model drop into chinese on long runs. I had to remind it to use english. Its still better than every gemma model I tried. Gemma deleted files on a harddrive to make space when there was over 2TB free. For long runs, gemma is useless.
- nolok 16d agoIf you care about points sure, but personnaly I care about price, performance, speed and reliability
- WarmWash 16d agoThe benchmark also doesn't include speed. You almost think something has gone wrong when using it because it returns full responses so incredibly fast.
- scrlk 16d agoNot just speed, also reliability. IME, Gemini's speed and quality doesn't degrade badly during weekday working hours compared to OAI, and especially Anthropic.
- ford 16d agoI've had Gemini model API use degrade the most out of OAI/Anthropic/Google (often "over capacity" vs true failures) Not sure on consumer/product use though
- scrlk 16d agoThat's interesting to hear. I should have added that I use Gemini through Google AI Studio as my general chat model, which probably explains our wildly different experiences.
- sotix 15d agoThis one uses that as a priority weight: https://winstonrc.github.io/ai-coding-agents-leaderboard/ https://winstonrc.github.io/ai-coding-agents-leaderboard/
- ttul 16d agoCrushing it on DeepSWE is a very big deal. Excited to give this a try.
- pietz 16d agoI know everyone is benchmaxxing but this one feels one step too far. Doesn't DeepSWE have both public and private tasks? I'd love to see the diff here. It looks more like Google execs losing their mind and pressuring researchers to put DeepSWE directly into the training set.
- re-thc 16d ago> DeepSWE is a very big deal It's clearly been "dealt with" already. When it launched we had interesting gaps and definitely differences. Now every new release is "crushing it".
- ttul 16d agoWill look forward to the "feel" of the model in real testing. But I agree that these benchmarks do get "dealt with" rapidly. That's a shame, but I guess it's the times we live in.
- rickdg 15d agoCheck DeepSWE for number of agent steps.
- satvikpendem 16d agoWe'll see about that. I suspect benchmaxxing as all the labs do as I haven't found Gemini models to be nearly as good in agentic engineering compared to Claude or GPT models.
- onlyrealcuzzo 16d agoAnd the benchmarks agreed with you... until now. So, yes, maybe it's still not - but this would be the only time it would be highly suspicious / obvious benchmaxxing / obviously bad benchmarks.
- NitpickLawyer 16d agoIf anything, gemini models are the least benchmaxxed out of any lab, IMO.
- onlyrealcuzzo 16d agoThe rumor is that 3.9 is an equal improvement in all directions, and that it should be another fast follow on like 3.7 and 3.8 were. It's almost across the board better than Terra at less than half the price. 3.9 is likely to approach Sol at the 1/10th the price. Hopefully OpenAI releases Astra first, and it's not only better than Sol but significantly cheaper, too.
- harmonic18374 16d agoCurious where did you hear this rumor?
- onlyrealcuzzo 16d agoAll the talk on Reddit on Gemini 3.8 discussions: https://www.reddit.com/search/?q=gemini+3.9&cId=1650e403-bcfa-402b-a352-9e6113d181e6&iId=f114a6dc-dd42-41ba-bf90-72d18852d52e https://www.reddit.com/search/?q=gemini+3.9&cId=1650e403-bcf...
- nolok 16d agoAccordit to reddit talk, Fable 5.1 is worse than Opus 4.6 and 8B models are smarter than Qwen 3.8 Max, I wouldn't take anything said there with any more reliability than an instagram short.
- unsupp0rted 15d agoReddit thinks Astra will be released today (Thursday/Friday)
- markasoftware 16d agoOn artificial analysis it's only equal to opus 5 medium effort. Opus 5 max scores 63. Further, opus 5 medium outputs 4x fewer tokens to achieve the same result, negating a lot of the speed difference.
- irishcoffee 16d agoA comparison to an artificial score and a comparison to “the same task” These folks must laugh themselves to sleep. This whole industry hoodwinked the masses. It’s impressive.
- wonnage 16d agoIt’s all just vibes
- theHocineSaad 16d agoAs of writing this comment, Claude Opus 5 has an intelligence score of 63, not 59 (it's not the same as Gemini 3.8 Flash). With a score of 59, Gemini 3.8 Flash is in eighth place, falling behind even Grok 4.6, Kimi k3, and GLM 5.3. https://imgur.com/a/BMOJBED https://imgur.com/a/BMOJBED
- Squarex 16d agoThey are all much larger and more expensive models. Google does not have a frontier model right now, but for cheap ones, they are better than event the chinese models now.
- pietz 16d agoThat's not being debated here. The initial reported numbers were false and this was simply pointed out. You're changing the subject.
- mattlondon 16d agoOpus 5 medium has the same score as 3.8 flash on artificial analysis intelligence index. Are you implying Google or Artificial Analysis are reporting false numbers? What's your source?
- asdfologist 16d agoBTW you're comparing 3.8 flash high to opus 5 medium. 3.8 flash medium scores lower.
- bertili 16d agoA fifth of the cost of Opus 5! Google is certainly pushing the completion with this.
- abirch 16d agoGemini hasn't failed me for personal usage yet. I haven't had the opportunity to use it at work.
- panarky 16d agoI've been using 3.7 Flash to audit the work of Opus High, and Flash finds lots of subtle and insidious defects even while all the unit tests are green. Then I tell Opus to read the audit report and implement what it agrees with. Flash is really good at this, and it is blazing fast in Antigravity CLI. Easily 10x faster than Opus. Can't wait to try 3.8 Flash. If it's good enough, maybe I'll switch Flash to primary and make Opus the auditor.
- porridgeraisin 16d agoYeah the speed in agy cli is amazing. Whole files get written and "py_compile"d in a single blink of the eye its crazy. In india, my telco gives me google ai pro for free. And agy with flash goes a long way.
- prodigycorp 16d agoit's very fast but it still doesnt come close to 5.6 sol, at least for me, in terms of gathering the context necessary to do extensive changes.
- MaxikCZ 16d agoIdk, was building/maintaining simple esp32 control program with antig/opus. After last update it defaulted to gflash3.7. I pasted an email requesting 2 changes into the chat prompt, it did one and took me 4 turns to get that one right.
- pkos98 16d agoWait a week with your judgement - most likely, Google is just bench-maxing very hard. If you look at the previous Flash models and the announcement on Google I/O, it was an absolute disaster. Reality diverged very much from the marketing (supposedly great benchmarks).
- jrflo 16d agosidenote, but wow sonnet 5 is shockingly bad on this benchmark.
- notatoad 16d agosonnet 5 is bad by almost any metric. anthropic really needs something to address the cheaper end of the market before they get left behind. Sonnet 5 sucks, and Haiku hasn't been updated in a year. meanwhile we've got gemini flash, luna, and GLM5.3 all delivering 90% of the performance for a small fraction of the cost. paying $25/mTok is going to start looking pretty silly soon.
- surgical_fire 15d ago> flash, luna, and GLM5.3 all delivering 90% of the performance for a small fraction of the cost. I find GLM5.3 so much better than Sonnet it is not even funny. Sonnet behaves like a cheap model while being very expensive.
- kimjune01 16d agodeepswe is public and can be considered contaminated.
- WhitneyLand 16d agoThere are important gaps in that hot take. For example, it's not even close to Opus 5 on Terminal-bench 4.0, 19.1% vs. 51.8%.
- wwind123 15d agoI've been trying this Gemini 3.8 Flash for a day. Looks not much different than Gemini 3.7 Flash in my use case: I have Codex (gpt-5.6 sol) write up a design plan to implement a feature or refactor a portion of a system I am building, and have Claude (Opus-5) and Gemini (3.8 Flash) review and critique the plan, until all problems are addressed by Codex and approved by the reviewers; then have a cheaper model of Codex (gpt-5.6 luna) implement the plan, and still have Claude (Opus-5) and Gemini (3.8 Flash) review and critique the implementation, until all problems are addressed by Codex and approved by the reviewers. The result is the same as the previous Gemini 3.6/3.7 Flash days: Claude could always note much more problems in Codex's plan and implementation than Gemini could - the ratio is like 10:1. I occasionally switch the roles between Codex and Claude, and result is the same, Codex could always catch much more problems in Claude's plan and implementation, than Gemini could. So I am guessing in a relatedly complex codebase, Gemini is much less effective in acting as a guardrail (or a senior engineer/team lead) than the other SOTA models.
- camkego 15d agoI find this very interesting, I wonder if there is a public benchmark that reflects this “red team coding critique” aspect of the current SOTA model that reflects what you have observed. It would be really useful to observe this in a benchmark vs. the more common “go implement this, or fix this bug” type benchmarks that seem to be prevalent.
- wwind123 15d agoYeah, my tool to automate these review loops is https://github.com/wwind123/coding-review-agent-loop https://github.com/wwind123/coding-review-agent-loop . It's basically a script calling Claude, Codex and Antigravity CLI's. The benefit of using CLI's is, the tool uses quota in your subscription plan of these AI providers, which is much cheaper than using extra tokens from the same providers to do the same thing. A couple of months ago (before opus-5 and gpt-5.6 sol), The ratio of problems caught by codex/claude vs gemini was more like 2:1 to 3:1. But now it seems codex and claude have made huge leaps and gemini is more or less staying put.
- rickdg 15d agoCheck number of agent steps.