6 ms·
They are all much larger and more expensive models. Google does not have a frontier model right now, but for cheap ones, they are better than event the chinese
by Squarex 16d ago
They are all much larger and more expensive models. Google does not have a frontier model right now, but for cheap ones, they are better than event the chinese models now.
- pietz 16d agoThat's not being debated here. The initial reported numbers were false and this was simply pointed out. You're changing the subject.
- mattlondon 16d agoOpus 5 medium has the same score as 3.8 flash on artificial analysis intelligence index. Are you implying Google or Artificial Analysis are reporting false numbers? What's your source?
- asdfologist 16d agoBTW you're comparing 3.8 flash high to opus 5 medium. 3.8 flash medium scores lower.
- WarmWash 16d agoFlash models are on the order of 1/10th the size of Opus models, so some flex in the thinking level is fair.
- Topfi 16d agoFlash is just a name with no defined or consistent meaning even within labs, let alone between them. Considering both are closed weight, there is no way to truly assess how big the size delta between the two is. Then again, who cares about size, performance and end-to-end speed+cost are what matters along with task adherence, task assessment and so on. Model size also can not be inferred by tokens/sec for a multitude of reasons, but to showcase two examples, Opus 5 and Sonnet 5, as well as Gemini 3.1 Pro Preview and 3.1 Flash have each very comparable output speeds when using the same deployment as a basis for comparison, despite it being very likely that within their generation, the former are larger than the latter. Feel the need to mention this, as I unfortunately stumble upon so many poorly reasoned, speculative hype post trying to infer model size via utterly unreliable metrics, not based in actual data. It’s like comments below arguing about the reasoning levels not normalized to some metric (like cost, output token amount or duration) but just the labels or high, max, medium, etc. Those mean almost nothing even when comparing models based on the same pretrain (just compare GPT-5.4 to GPT-5.2), they mean less than nothing comparing different labs releases.
- WarmWash 16d agoIt's not totally a mystery https://arxiv.org/html/2604.24827v1 https://arxiv.org/html/2604.24827v1 The short of it is by using hard facts knowledge that is difficult to compress, and then quizzing models on these facts and calibrating against a bunch of open models, you can kind of feel out the size of closed models.
- Topfi 16d agoI really like that one, but it kinda highlights what I could have far better explained. Their 90% PI is three times in both directions. Between 3T and 24T for GPT-5.5. That’s a massively wide, inaccurate and at best barely informative range, demonstrating that even the most well thought out method will yield little usable information. Additionally, I got some private evaluation taking a similar approach towards gauging models in topics I’ve found either over or underfitted by labs. If we just used that to rank models (not get a potential size range but just a rough order) Thinking Machines Inkling would need to be lager than Fable 5.
- nomel 16d agoWhen comparing closed models, the only thing that actually matters to anyone using them is some mix of cost and speed. Considering how much memory a server is using, when evaluating models that you'll never have access to in order to host yourself, doesn't really make sense.
- imtringued 16d agoYour comment is really strange, why are you defensive towards WarmWash when gemini flash 3.8 high is both 6 times faster and costs less, while having the same intelligence score as claude opus 5 medium? >Considering how much memory a server is using, when evaluating models that you'll never have access to in order to host yourself, doesn't really make sense. This entire sentence makes no sense given what is being discussed. https://artificialanalysis.ai/models/gemini-3-8-flash https://artificialanalysis.ai/models/gemini-3-8-flash https://artificialanalysis.ai/models/claude-opus-5-medium https://artificialanalysis.ai/models/claude-opus-5-medium
- nomel 14d agoI was being pragmatic. These are closed models on closed systems that you cannot hope to host. They are only available as black boxes available over web APIs served by their owners. Within that black box perspective, that we're force to have, the size of the model is, quite literally, just how much memory that server is using. intelligence/model size is not a useful metric for a black box user. intelligence/cost and intelligence/speed is a useful metric for a black box user. Yes, it's cool, but as a black box user, the amount of memory a model is using on a server that I do not own has exactly zero practical use to me. Cheers!
- duplessitous 16d ago> [...] shows an intelligence score of 59, the same as Opus 5 medium! Nothing here is false, you are simply confused. You either didn't read what they wrote in its entirety or decided to reinterpret what they did write.
- knollimar 16d ago"Beating opus" is the false part, no?
- imtringued 16d agoStop lying. mattlondon said "gemini-3-8-flash shows an intelligence score of 59" which is undeniably correct. You can't say that number is false. You're literally lying. All you had to do is go hover your mouse over "Models" in the top bar, hover over Claude Opus 5 and and click on medium: https://imgur.com/mlRCrt1 https://imgur.com/mlRCrt1 When you do that you arrive on this page: https://artificialanalysis.ai/models/claude-opus-5-medium https://artificialanalysis.ai/models/claude-opus-5-medium The gemini flash page for reference: https://artificialanalysis.ai/models/gemini-3-8-flash https://artificialanalysis.ai/models/gemini-3-8-flash You have to be an incredibly dishonest person to see a 59 on both pages and say "the initial reported numbers were false and this was simply pointed out. You're changing the subject".
- porphyra 16d agoBetter than even the Chinese models? That's a difficult-to-quantify, extremely rapidly moving target. Just today, Qwen 3.8 Max 0902 came out with a huge improvement over the previous Qwen 3.8 Max.
- spwa4 16d ago> "Google does not have a frontier model right now, but for cheap ones, they are better than event the chinese models now." Just wow. Someone actually said this.
- zeckalpha 15d agoGoogle is targeting a different segment of the frontier.