16 ms·
Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
https://console.cloud.google.com/agent-platform/publishers/google/model-garden/gemini-3.6-flash https://console.cloud.google.com/agent-platform/publishers/g...
- yanis_t 2mo agoThe benchmarks are not particularly impressive. I suppose they needed to release something since the long pause. But not clear why would I use it now.
- dumberquestions 2mo ago"..and in some benchmarks like DeepSWE by Datacurve, we observe up to 65%, all at a lower cost per output token." "3.6 Flash delivers higher precision with fewer unwanted code edits and reduced execution loops, as seen in DeepSWE (49% vs. 37%)" So which one is it? 65% or 49%?
- petu 2mo agoFirst sentence is about token efficiency.
- dumberquestions 2mo agoYou're right, should've gotten some LLM to summarize it instead of skimming.
- semilin 2mo agoOr you could have read it more closely before posting a comment saying it didn't make sense. You know, the old school way.
- dumberquestions 2mo agoIf I'm going to read it wrong might as well have an LLM to blame.
- jgbuddy 2mo agoIt is both less intelligent and more expensive than GLM-5.2, while being closed weight.
- drob518 2mo agoBut they make up for it by shipping it late.
- SwellJoe 2mo agoIt's also got vision and audio. So, the better comparison is any of the other large Chinese open models that are better and cheaper than Gemini Flash.
- lenerdenator 2mo agoIt'd be interesting to know how much the Intelligence as a Service angle serves as a value-add in the minds of Google's executives. You can get decent open-weight models now. That's not difficult. The difficulty is 1) running them and 2) compliance. My company runs Claude on GCP's Vertex AI solution. We're in the US healthcare IT space, so the models need to be from somewhere that American healthcare agencies and companies have traditionally been okay with sourcing code from - which means the US, Canada, and maybe Europe. The stuff that handles PHI/PII must be in the US. The expense of hosting is more of a PITA than most customers want to go through this early in the technology's lifecycle, and intelligence gains are simply a matter of degree for most business tasks. In theory, we could find some open-weight model (likely from China) for our development agentic work and host it anywhere you can host AI models. We don't, though, and I think Google, OpenAI/Microsoft, and Anthropic see that as the core of their business.
- killix 2mo ago[flagged]
- dyauspitr 2mo agoIt’s multimodal though.
- metalliqaz 2mo agoOther discussion from a few minutes earlier: https://news.ycombinator.com/item?id=48993130 https://news.ycombinator.com/item?id=48993130
- velominati 2mo agoWow - Google does not even bother to show benchmarks of these models compared to the frontier and Chinese labs - only against previous versions. I'm not surprised. Having worked there for years it was amazing just how inwardly looking the company is.
- dvduval 2mo agoIt does seem like their releases are getting closer together. I get the feeling they realized they were trying to roll out to their entire ecosystem and now they’re focusing more just directly on the AI model itself. I think give it a little time and they’ll start to be one of the competitors too.
- singingtoday 2mo agoI'm more excited for 3.5 pro. Gemini has fallen behind in some areas, but is still one of the best multimodal models. Has anybody found any models better at image or audio analysis?
- ianhawes 2mo agoCame here to ask basically this. We use 3.1 Pro internally and it's great.
- npn 2mo agotested the models on aistudio. despite that the knowledge cut off is march 2026 it still knows nothing about 2025! you can check by asking "list notable world events in 2025, only list unplanned" on aistudio. or you can ask for Charlie Kirk, it also does not know. I tried it multiple time to ensure that I didn't not get routed to older models! > but google has search irrelevant, without deeper knowledge about cutting edge technologies or latest libraries, all of it suggestions are crap. even you ask it to search it will still use outdated keyword thus only getting outdated information. in other word, what a disaster!
- npn 2mo agoupdate: the knowledge cut off date is "unknown" now. funny because some people downvoted me believed that there is no relation between knowledge cut off date and real world events. that's not how it works!
- nsbk 2mo agoIt is 17% more token-efficient than 3.5 and performs significantly better in coding and tool usage benchmarks. It is also cheaper than 3.5: > This enhanced efficiency is also combined with a lower price than 3.5 Flash. At $1.50/1M input tokens and $7.50/1M output tokens, 3.6 Flash reduces the overall cost per agentic task, making agents more cost-effective to build and run.
- ilreb 2mo agodupe? https://news.ycombinator.com/item?id=48993130 https://news.ycombinator.com/item?id=48993130
- m_w_ 2mo agoIt's a bit disheartening to see no comparison to other models here - and I'm not sure this pushes the curve anywhere. 3.6 flash is more expensive than GLM 5.2 - but seemingly worse, although this post is really light (lite?) on details. It seemed for a time that Google had finally gotten the ball rolling, but I'm doubting that more and more as time passes. We'll see what happens with 3.5 pro I suppose.
- dd8601fn 2mo ago> really light (lite?) on Light. Lite is product marketing seepage.
- crab_galaxy 2mo agoYeah that’s the joke :p
- dd8601fn 2mo agoSorry, it went right over my head!
- kzrdude 2mo agoI've never questioned the word lite before because it's existed my whole life.. So does it make sense? Why does it exist and where does it come from? More than coming from "light".
- Alpha3031 2mo agoThere's an Merriam-Webster article on it: https://www.merriam-webster.com/wordplay/lite-word-history https://www.merriam-webster.com/wordplay/lite-word-history
- lopatin 2mo ago[dead]
- armarr 2mo ago
- b473a 2mo agoNo word about updating Jules, which is still stuck on 3.1 Pro. I get that it's probably niche but I've really appreciated basically being able to give directions to Jules on my phone, then reviewing and merging a GitHub PR fifteen minutes later. It's been great for getting some progress in on a few personal projects during my commute when I can't exactly pull out my laptop. Anyone have any good alternatives?
- haberdasher 2mo agoClaude Code
- steven_pareto 2mo agoIf you own a Raspberry Pi or similar: Hermes + Tailscale + iSH over tmux.
- deleted 2mo ago[deleted]
- aweb 2mo agoBoth Claude and Codex can code in the cloud, it works quite well! I tested Jules and while the idea is good in theory, I found the model's intelligence to be very lackluster.
- b473a 2mo agoShame. I'm on the $20/mo Gemini Pro plan because the 5tb of cloud storage and the youtube premium lite were good enough perks, and my coding complexity needs were light enough for me to overlook Claude or Codex. But Antigravity is working better than Jules and it's basically giving me a taste of what I'm missing and it's harder to justify not trying out the competitors.
- christoff12 2mo agoI have no affiliations with the team or product, but Superconductor reminded me of Jules when I tried it a couple of months ago. It might be overkill features-wise, but there's a free tier and it likely won't be left for dead anytime soon.
- geooff_ 2mo agoAt this point just put the Pareto in the bag bruh
- ConfusedDog 2mo agoWhy would 3.6 flash perform a little worse than 3.5 flash on Artificial Analysis Coding Index... https://artificialanalysis.ai/models/gemini-3-6-flash?intelligence=coding-index https://artificialanalysis.ai/models/gemini-3-6-flash?intell...
- sosodev 2mo agoBecause AA Coding "Index" consists only of two benchmarks (Terminal-Bench v2.1, SciCode) and generally fails to be meaningfully representative of agentic coding capabilities.
- Alifatisk 2mo agoWhats a better option for AA Coding Index?
- WASDx 2mo agoDeepSWE and FrontierCode are more realistic if you read up on what they actually measure. But the most realistic is to try it yourself. Benchmarks can only vaguely represent typical usage, and how you judge the result. Giving the same real task you have to a few models will make you understand them better than chasing benchmarks.
- kimjune01 2mo agoit would be nice if these benchmark reports actually specified which tasks they passed and which ones they didn't.
- firethunder7 2mo agoAA coding index has been updated to use DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA.
- sosodev 2mo agoWhen? It literally says on the page for Gemini 3.6 Flash "Artificial Analysis Coding Index represents the weighted average of coding benchmarks in the Artificial Analysis Intelligence Index (Terminal-Bench v2.1, SciCode)"
- dankai 2mo agoUnfortunately says more about how competitive 3.5 pro would be today at the frontier if they forgo it for 3.6 flash.
- catigula 2mo ago"We made 3.6/4 Pro, but it sucks, so this is the distilled model" vibes.
- drob518 2mo agoThat’s the fear.
- mfkrause 2mo agoPretty underwhelming, as expected honestly. I don't want to know what morale is like at DeepMind right now.
- drob518 2mo agoYep, agreed. They still are not releasing anything frontier-class (Gemini Pro) at this point. Feels to me that they keep getting scooped by others (e.g. Kimi 3) and then are retrenching.
- WarmWash 2mo agoEspecially when Google owns 15% of anthropic and serves them compute. Double especially when your boss (Hassibis) is also an early investor in Anthropic. Hell his NW might be more Anthropic than Google.
- kilroy123 2mo agoI deeply wish Google would focus on models like Gemma. Small, powerful, open-weight models you can run on phones or regular computer hardware.
- lanthissa 2mo agoi mean eventually then will, losing means open source, vertically integrated hardware means you can opensource and win on cost
- mediaman 2mo agoGemma 4 was released in April. It's a good series of multimodal models.
- accountrequired 2mo agogemma 4 thinks joe biden is president
- mediaman 2mo agoSmall open source models shouldn't be used for world knowledge, that's not their purpose.
- Petersipoi 2mo agoWhy not? Seems like a cop out. Being able to ask questions to small open models seems.... obviously useful?
- mediaman 2mo agoBecause they don't have a lot of parameters to store general Wikipedia knowledge. They're small. Use big models that have high parameter capacity to store general information. Or build a harness around the small model that searches a knowledge base/internet. Use the right tool for the job. It's like asking why a screwdriver isn't good at sawing wood, or calling C a terrible language because it's hard to make CRUD apps with it.
- doctoboggan 2mo agoI have a side business selling custom fingerprint jewelry and I use gemini nano banana to clean up customer submitted fingerprint images. This was a step I used to do by hand at 10 - 15 minutes per image and nano banana is the first model that is able to do the task (it is astonishingly good at it). I can't wait to see what the next nano banana can do, hopefully its released soon.
- cube00 2mo agoAre your customers clearly informed that you're sending their immutable fingerprints to an AI service?
- poisonborz 2mo agoYes this is extremely unresponsible if so. Fingerprints are legally protected biometric data in most juristictions.
- HDBaseT 2mo agoIt is not mentioned in their FAQ. [0]. [0] https://lulimjewelry.com/pages/faq https://lulimjewelry.com/pages/faq
- dinkelberg 2mo agoTheir shop is linked to in the bio. They don't seem to have a privacy policy up on the site. When in the checkout form it links to the generic Shopify privacy policy. No hints to the fact that uploaded images are processed by third parties, as far as I can tell. That should be corrected for sure.
- pietz 2mo agoAre they comparing 3.6 Flash to 5.6 Luna and losing? That's ruff.
- polski-g 2mo agoWhy wouldn't they? Luna isn't a Flash model. OpenAI hasn't released a flash-equivalent model since gpt-oss-120b.
- deleted 2mo ago[deleted]
- pietz 2mo agoDid you ask me a question and then answered it yourself in the very next sentence? Anyway, given that both Gemini and OpenAI have 3 sizes of models, one would think Google compares their medium size to OpenAIs.
- primaprashant 2mo agoPricing per million input/output tokens: 2.5 Flash: $0.3 / $2.5 3.0 Flash: $0.5 / $3 3.5 Flash: $1.5 / $9 3.6 Flash: $1.5 / $7.5 --- 2.5 Flash-Lite: $0.1 / $0.4 3.1 Flash-Lite: $0.25 / $1.5 3.5 Flash-Lite: $0.3 / $2.5
- jjice 2mo agoAm I off, or does Google have the pricing that varies the most between model generation releases?
- m_w_ 2mo agoIt seems that they're trying to push up-market, or at least they were. Given the extremely competitive releases of GLM 5.2 and DeepSeek V4 (both pro and flash), I don't think there'll be appetite for it.
- urbsgpw 2mo agoIt seems like they're sticking to a static pricing plan that was made when the only relevant competition were the US labs (im not counting deepseek 2025 as serious competition -> glm and then kimi on the other hand, now that's a different story).
- LaurensBER 2mo agoPricing often reflects what the vendors (expects) the customer is willing to pay. It seems that Google is still trying to find their niche in the market.
- mchusma 2mo ago3.6 Flash would be a great model at 3.0 flash pricing. At this pricing, its thoroughly trounced by about 10 models on cost/performance including Grok 4.5. 3.5 Flash-ite would be a great model at 2.5 flash-lite pricing, as is, its trounced by many models including Deepseek v4 Flash. As is, they are thoroughly outclassed for most usecases. I will say the one area where i do see Gemini punching above its weight class is in tasks that are effectively "Google this for me" / knowledge stuff. So it does have a role, and I do use it. So while I think Google is still in a strong position overall, they are really stuck as a tier 2 AI player right now with text models. They are tier 1 in bio, images, and video.
- holistio 2mo agoThey are comparing against their own previous models instead of competitors. Not a great sign.
- ChrisArchitect 2mo agoSome more discussion: Gemini 3.6 Flash https://news.ycombinator.com/item?id=48993130 https://news.ycombinator.com/item?id=48993130
- ComputerGuru 2mo agoSo 3.6 Flash is a somewhat of an admission that Google miscalculated by charging 3-5x for 3.5 Flash what it did for 3.0 Flash (3x input and output costs plus large token inefficiency changes) despite only modest improvements? 3.5 Flash Lite is only a hair cheaper than 3.0 Flash, but I think 3.0 Flash is a massively more capable model?
- tiahura 2mo ago3.5 Pro must really suck.
- WarmWash 2mo agoThe mention of an "ambitious" gemini 4 pre-train signals to me that 3.5 pro is probably a lost cause. That being said, it seems that Gemini is still the best image analysis model, so hopefully 3.6 flash builds on this even more.
- canergl 2mo ago2 red flags 1- no comparison with gemini 3.1 pro 2- no comparison with any other model
- gs17 2mo agoThe model card has comparisons with both 3.1 Pro and other models: https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-6-Flash-Model-Card.pdf https://storage.googleapis.com/deepmind-media/Model-Cards/Ge...
- youssefarizk 2mo ago3.5-lite is the real showpiece here; agentic models of this size are a huge value-add for 90% of knowledge work agent tasks
- sreekanth850 2mo agoGoogle is walking backwards, with such a pile of cash in pocket, i feel they are doomed.
- thevinter 2mo agoI struggle to see any value in this when DeepSeek is still a thing.
- anthonypasq 2mo agomultimodal + latency
- kzrdude 2mo agoIt's smarter than DeepSeek v4 Pro (preview) on several benchmarks, like HLE.
- WhitneyLand 2mo agoThe silence is deafening. Google watches over the last few months a flat out assault on the Pareto curve from American and Chinese companies. Release after release pushing the boundaries of frontier intelligence and price/performance. And the response from arguably the biggest AI research labs in the world by headcount is Flash 3.6. What do you do when you are given essentially unlimited resources and still find yourself falling behind?
- kthinckley 2mo agoGoogle desperately needs to make some leadership changes within their Gemini team now that they've been surpassed by 3-5 open weight models and risk loosing frontier status all together in the near future.
- alephnerd 2mo agoOpen weight models aren't likely to be open weight in the long-term. China has started considering export controlling and limiting access to model weights [0]. [0] - https://www.ft.com/content/6049a031-9e9b-464c-97bb-414da04d5a6a https://www.ft.com/content/6049a031-9e9b-464c-97bb-414da04d5...
- ErneX 2mo agoThat contradicts this: https://www.wsj.com/tech/ai/chinas-xi-touts-open-source-ai-and-takes-a-swipe-at-u-s-dominance-1eaa5cfe https://www.wsj.com/tech/ai/chinas-xi-touts-open-source-ai-a... So who even knows.
- logicchains 2mo agoThey don't need to be open weight in the long term; once there's an open-weight Fable-level model with 1M context it'll be pretty much good enough for all coding tasks, no need for new models.
- llmslave 2mo agoI keep saying this and people dont believe me, but I have b2b saas systems with actual agents running around the clock, and the performance/stability of the flash model is higher than most other models. Meaning, its predictable with tool calls, wont spin off a million tools/do weird behavior, its reasonable. Even sonnet in a real world decision making scenario is not reliable, or will reason so long its incredibly expensive. The benchmarks arent catching all the value, and most people have never actually ran an ai agent in a real context that matters
- sureMan6 2mo agoWho's most people? What are you talking about? Most people here use agents every day and I wouldn't trust flash or pro to touch any important project of mine because they're both terrible compared to the competition, waste of time every time I give them a chance
- llmslave 2mo agoI mean like an ai agent doing some sort of HR work, not a coding agent. Very few businesses are trusting an autonomous agent.
- onlyrealcuzzo 2mo agoGemini 3.5 flash is already a pretty good model. But, unfortunately, the primary way you can interact with it for coding is through Antigravity - which is actively developer hostile. It doesn't matter how good the model is if you're (mostly) forced to use it in Antigravity - which turns any model into crap. Wake me up when Antigravity doesn't suck.
- postalcoder 2mo agoI wonder how big the Pro model is that Google is using behind the scenes to train these smaller ones. Going on baseless speculation, the lack of accompanying pro models with these flash releases either means: 1) the model is too big to be economical, 2) google doesn't have the compute to serve the big model, 3) their big model has too many alignment issues to serve to the public. edit: looks like benchmarks are up on https://artificialanalysis.ai/models/gemini-3-6-flash https://artificialanalysis.ai/models/gemini-3-6-flash. It's solidly middle-of-pack. However, if you want to be most fair to flash, look at the intelligence vs time per task and intelligence vs outputspeed benchmarks. This is a very fast model. edit 2: I use antigravity from time to time and in my experience, 3.5 flash is an underrated model, so long as you know what it's good for. It's very good at frontend (much better than gpt 5.5) and it's fast, so it's a great tool for iteration. I expect 3.6 to be no different.
- anthonypasq 2mo agoLogan Kilpatrick said on an interview not too long ago that flash 3 and 3.5 are the same pre-train. all gains on top of 3 flash are post-training
- mchusma 2mo agoMaybe, but they said they have “started” the Gemini 4 pretrain. So not having done any significant pretrain in a year or so seems odd to me.
- WarmWash 2mo agoPre-trains take a huge chunk of your compute offline, incurring both an raw expense (24/7 max power for all training clusters) and an opportunity cost (could have sold excess compute during that time). They also don't come with any great guarantees, as lots of techniques look good on small scale and crumble or plateau once scaled.
- petercooper 2mo agoI wonder if the broad use of AI overviews on Google search results is having an impact. Maybe the numbers make it more profitable to use their compute on several billion searches a day rather than selling API access.
- parsimo2010 2mo agoFeels like they released this to ride the wave of press of GPT-5.6, Kimi K3, and Qwen 3.8. Doesn't feel like Google has much substance with this post except a bump in version and tweaked their pricing.
- primaprashant 2mo agoA couple tidbits: > Beyond today’s releases, Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it’s ready. > We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress.
- cubefox 2mo agoI guess they meant to release Gemini 3.5 Pro shortly after 3.5 Flash, but then Mythos/Fable and later GPT-5.6 came out with higher performance than 3.5 Pro, so the managers decided not to release it.
- jdiff 2mo agoThat reasoning didn't stop them from releasing this batch of models, admittedly that may be less face to lose away from the flagship position.
- kingstnap 2mo agoWhile them fixing token bloat on 3.5 Flash is good work. That paragraph was the real highlight. Hopefully 3.5 Pro is soon, and that Gemini 4 can be here end of year and finally have an updated knowledge cutoff.
- deleted 2mo ago[deleted]
- deleted 2mo ago[deleted]
- prox 2mo agoWhat is a “knowledge cutoff” ? —Ah, got it, it knows more about recent times.
- tulio_ribeiro 2mo agoI’d just like to add, and this may interest you both, that I’m not disagreeing with either of you. When 3.5 Flash first showed up in AI Studio, its model card said it had a March 2026 knowledge cutoff. After the backlash over it apparently knowing nothing past December 2024, the card was changed to read: “Knowledge cutoff: Unknown.” Maybe the timing was coincidental, but I doubt it.
- zb3 2mo ago> we have taken an intentional approach to deploying 3.5 Flash Cyber. The model will be exclusively available to governments and trusted partners Screw your government! US and Israeli governments should get the least access, but of course we all know they'll be the (only) ones to get full unfiltered access.
- m4tthumphrey 2mo agoI'm going to get downvoted/flagged but I feel like we need a new type of "Show HN/Tell HN" etc for "New AI Model Available". Front page is tedious these days.
- tremarley 2mo agoIf you refresh the Home page, once a day. The top post will likely be 'New AI Model Available '
- Alifatisk 2mo ago> I'm going to get downvoted/flagged but [...] Why do you care that much? Just say it. You're letting an imaginary score determine if you should express your suggestion for improvement. That's a bit wild.
- m4tthumphrey 2mo agoNo I'm not, I expressed it... I prefaced it with that to simply say I knew it'd be an unpopular opinion.
- ece 2mo agoJust switched to AI Plus from Pro, seems like I won't be missing much.
- spyckie2 2mo agoGoogle seems to have anorexia when it comes to model intelligence. They have an internal hard constraint on price per token it seems, and they are trying to squeeze out intelligence with limited compute. I wonder if there is something with their TPU cycles that makes them want to postpone training a new model. My guess is that they have been on the same base model for 6 months and they may have waited for the next gen TPUs to train Gemini 4, which greatly limits how much intelligence they can increase and forces them to do cost efficiency increases.
- WarmWash 2mo agoGoogle Cloud is probably Google Deepminds biggest competitor. Big company kinda bullshit.
- platinumrad 2mo agoHow so?
- WarmWash 2mo agoGoogle cloud sells compute out from under Deepmind to other labs. So they basically are in competition with Google cloud for compute.
- logicchains 2mo agoI'd guess they did model-hardware codesign but the design ended up limiting the scaling capability of the model (i.e. they overoptimized too soon).
- JacobAsmuth 2mo agoCould it be that they have to serve their models to billions of users?
- ur-whale 2mo ago> Could it be that they have to serve their models to billions of users? And how is that different from their competitors exactly?
- AussieWog93 2mo agoA lot of disappointment here in the comments, but models like these aren't meant to compete with the likes of Fable or GPT 5.6. I use 3.1 Flash Lite regularly to classify listings on eCommerce websites. It's great for this task - fast, cheap and accurate. In fact, it was the single best model we tried in terms of the speed vs accuracy vs price tradeoffs - including the Chinese models. Of course, 3.5 Flash was more accurate but the 5x cost increase couldn't be justified. 3.5 Flash Lite sounds like it could be a strict upgrade for our use case, without a significant increase in costs or drop in speed. It's not GPT-6 but it's not trying to be. It's a completely different tool and great at what it does.
- fur-tea-laser 2mo agonot a google fanboy by any stretch... though i've been thrilled with the flash line of models... i exclusively use it on high, and have found it to be a great fit for increasing productivity 10-fold while maintaining quality... sure it can't just go off and one-shot a bunch of work, but at the complexity level i tend to work at, neither can the frontier in a robust way that i can be confident in... sure i have to be in the loop more, but that helps keep me grounded and course-correct earlier before wasting tokens... and when you sufficiently spec out a coding/software problem, and i mean really document all of the critical nuance, it will successfully satisfy the constraints... the quality is rarely acceptable on first-pass, but it forces me to stay connected to the architecture more than i would be if using a frontier model... i've found this to be a happy middle-ground of productivity and awareness...
- swe_dima 2mo agoIt's scary relying on Google's models. I have a very price sensitive workload that used to run on flash 2.5 lite - it's deprecated now. The replacement 3.1 flash lite is a lot more expensive, but now also has a sunset date. 3.5 flash lite is even more expensive. So the price is rising and you have no choice but to keep paying more and more.
- zuzululu 2mo agosame I just switched to OpenAI after using flash 2.5 lite for almost everything at our company. We spent thousands just to build this workflow now Google says screw off
- rayboy1995 2mo agoI moved directly from 2.5 flash lite to deepseek v4 flash, its already cheaper and if your prompt caching is good you can save so much more money.
- binary132 2mo agocould you explain how to optimize prompt caching or point to a doc about it?
- NeutralForest 2mo agoAnything Sam Rose is worth reading: https://ngrok.com/blog/prompt-caching https://ngrok.com/blog/prompt-caching but the implementation will be up to your provider and harness, for deepseek, they expose some numbers: https://api-docs.deepseek.com/guides/kv_cache/ https://api-docs.deepseek.com/guides/kv_cache/ and Anthropic has a list of actions invalidating your cache: https://platform.claude.com/docs/en/build-with-claude/prompt-caching#what-invalidates-the-cache https://platform.claude.com/docs/en/build-with-claude/prompt... Basically, you avoid anything dynamic: model change, tool change, etc it's also important that your system prompt or main prompt doesn't have non-static data like the date/time/place or someone's name (the person you interact with in a chatbot for example). That should be left to tool call or search.
- simonw 2mo agoPelicans for 3.6 Flash and 3.5 Flash-Lite (Cyber isn't available to me through the API yet.) https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fraw.githubusercontent.com%2Fsimonw%2Fllm-gemini%2Frefs%2Fheads%2Fclaude%2Fsession-4d1s16%2Flogs.md https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
- rednb 2mo agoI am growing tired of these pelicans posts every time a new model is published. Feels to me like low effort personal brand promotion. Just sharing my 2 cents.
- tomrod 2mo agoIts a nice benchmark. Like hearing the ice cream truck on a summer day.
- risyachka 2mo agoAt this point it does not show anything as models are fine tuned on all kinds of benchmarks.
- SoMomentary 2mo agoI thought the Gemini 3.5 Flash Lite response was quite telling myself. I personally like the Pelican SVG test, to me it is still a charming snapshot of model performance anecdata. No one would argue it's rigorous but I don't think it was ever intended to be. I get people burning out on the pelican SVG test alongside the rest of the AI burnout, but I guess for myself I'm just choosing to keep enjoying it while I still can.
- hypfer 2mo agoIt's both. I agree to rednb that at this point it feels like rather obvious brand building, but also, I agree with you that some value is in it. It does not feel all that authentic though, and it's good to react allergically to lack of authenticity. Bad for a lot of business models, but good for humanity.
- lambda 2mo ago3.6 Flash scores exactly the same as 3.5 Flash on the Artificial Analysis index. Better on some tasks, worse on others. Mostly within what I'd consider the noise window. Looks pretty much indistinguishable from 3.5 Flash, at least on these benchmarks: https://artificialanalysis.ai/models/gemini-3-6-flash https://artificialanalysis.ai/models/gemini-3-6-flash
- theplumber 2mo agoI think it’s safe to say Google seems a bit out of the top AI competition now. The “cyber” stuff also starts to become laughable with open models providing the full power without the crap Anthropic, Google, OpenAI are trying to frontload on you(I.e you are not allowed to develop/review a login system, pay a special cyber operation team to do it for you). They really deserve to become irrelevant in the future of AI.
- metahost 2mo agoSo about the same “intelligence” as Muse Spark 1.1 but 2x faster and about 2x as expensive.
- stonewhite 2mo agoGoogle somehow managed to snatch defeat from the jaws of success with their AI products. They literally forced me and my company out of Antigravity by phasing out AI Ultra subscription without any proper product follow-up. Antigravity IDE cannot even have poweruser subscriptions now from Google Workspace an Gemini Enterprise Agent Platform cannot be attached to Antigravity IDE. Gemini Enterprise Agent Platform has an incredibly abysmal setup process, and if I want to limit spending per-user I have to create projects per user. The fact that you cannot activate Anthropic models on it if the billing still has free credits is almost a joke. I was a big proponent of Google and Gemini, but they left us reeling with their abrupt product decisions. Forced us to buy $200 subscriptions directly from Anthropic/OpenAI.
- ikiris 2mo agoSnatching defeat from the jaws of victory is the specialty of product managers.
- ngrilly 2mo agoAlso a big proponent of Google and Gemini, but their stubbornness in artificially splitting their consumer and enterprise products is extremely annoying. It's pretty weird that I have access to more powerful tools when using my personal Google account compared to my corporate Google Workspace account.
- vel0city 2mo agoAs someone who's been using Workspace as a personal email account for over a decade this has been such a struggle forever. Just lots of odd limitations to feature sets all over the place. When they swapped Google Assistant for Gemini as the default voice provider in Android Auto it was so annoying. My wife's non-work space account can get Gemini to do the normal things like play music and what not, but my Workspace one can't do much of anything at all. I can talk about nearly any random topic with it, but getting it to change the playlist, nah, can't help you there. It's no surprise to me to see them fumble actually supporting a lot of the consumer features of Gemini into Workspace.
- ansuman441 2mo agoSpecific to task these can be huge plus point.
- dismalaf 2mo agoWith all the naysayers on Gemini models I'm curious how many people actually use Gemini regularly? For me, Gemini models are the most usable. Claude Opus and Mistral always try to turn queries into one-shot enormous commits, which just burns tokens, time and annoys me for something which is still wrong more often than not. Gemini seems far better at listening to instructions and giving me what I actually want, on top of using far fewer tokens and wasting my time. Fable is the only model that's come close to Gemini Pro for me. And as this is about Flash, it's exciting, I find Flash can usually get the right answer pretty quickly and without too much nonsense.
- dudeinhawaii 2mo agoI use all of the major providers daily and I tend to go to Gemini for "fast lookups" where a good enough answer is probably OK. I use ChatGPT and Claude for anything where it matters and generally when I invoke all three -- Gemini is the most surface level with responses, and also sycophantic. It gets worse from there. Gemini is terrible at agentic coding, primarily because Agy is terrible. I noticed Google updated Agy with this release, so perhaps that's finally going in a good direction. I'll have to test it. Thus far, my experience in countless experiments has been Gemini models being 2x faster yet with less depth in their solutions and a lot more going off track. I very rarely have to stop Codex or Claude Code sessions because they're doing something random and unexpected (or not asked for). I genuinely think Gemini models are brilliant but virtually useless in agentic scenarios in my experience of the last few years (2.5, 3, 3.1, 3.5). I should also note that Gemini web UI annoying resets to its lowest intelligence which feels scummy and Google is not transparent about what "extended thinking" really is. Past posts have pointed to "extended" being medium. Every other provider gives you the raw value (medium/high/etc). So honestly, I feel Google would rather I don't use their models. They just want to get a little bit of mindshare and stay in the conversation. I had the Ultra plan and cancelled it once it was apparent they were not improving the agentic experience nor trying to compete.
- tobias2014 2mo agoI agree, I see how Gemini itself with a usable harness can be excellent. But in agy with forced eager compaction (~125k with 3.1-pro, ~200k with flash) a kind of laziness and forgetting shows through that leads to an endless sequence of stopgap instructions, even with rigorous GEMINI.md and isolated task delegation and a good task tracking system. Agy is basically useless for more complex problems as far as I am concerned, at least when used somewhat autonomously as one could expect from claude. For strictly mechanical one-shot tasks it might be fine. I've spent way too much time working around these limitations instead of just continuing to use claude. Hoping that things would have improved with flash 3.6 I feel that it's actually worse in following instructions, and always acts even when just asked a question. If just agy offered a better experience and got rid of the terrible forced automatic eager compaction. PS: That opus-4.6 via agy works so much better points in another direction though!
- xnx 2mo agoProof-of-life release while they figure out how to have a competitive frontier model release. My hunch is they pushed too far in the "omni" model direction, that they made something so ungainly, it wasn't as good for normal tasks.
- Gecko4072 2mo agoI read this as a soft let down to not expect too much from 3.5 Pro. > We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress.
- zuzululu 2mo agoGoogle seems to be falling way behind the pack. antigravity cli is pure trash. gpt 3.5 pro is now behind and isn't released yet. GPT 6 and Fable 6 releasing next month. What the hell is going on over there ?
- zarzavat 2mo ago> What the hell is going on over there Google was late to coding agents and as-per-usual fucked it up with their crazy project management culture. Usually Google gets away with it due to inertia, however this time they are paying a heavy price because they missed out on the training data that Anthropic and OpenAI have gathered with claude and codex.
- cbossman 2mo ago[flagged]
- accountrequired 2mo agowhatever, dude. give gemma5
- revolvingthrow 2mo agoTons of guardrails, lazy model, super confusing plans, expensive 3.5/3.6 flash and lite and 3.5 pro MiA? Rough patch for google ai
- gabriel-uribe 2mo agoHaven't been excited for a Gemini release since December. Wild to see.
- XCSme 2mo agoI was expecting 3.6 Pro. It's been so long since the last Pro model...
- thebigspacefuck 2mo agoThey are working on coming up with a better code name. You know, something like ”Fable” or ”Sol”, gotta have one these days. Personally I think they should go with “Mafia”. How cool would that sound? 3.6 Mafia.
- dpacmittal 2mo ago3.6 Gangsta Pro
- baalimago 2mo agoNot good enough for high-end, not cheap enough to be for low-end. Next!
- u1hcw9nx 2mo agoGoogle has not changed. Following two facts are like tautologies by now. 1. Their AI efforts are very fundamental research oriented. They are really good at it. 2. Their productization sucks. The end products gets little attention compared to competition. It can be canceled at any time. You should never build anything around Google only APIs, AI or not.
- mythz 2mo agoAlways happy to see new Gemini releases as IMO Antigravity Pro 16.67/mo plan (Annual) is still the best plan available and have been pretty happy with Antigravity IDE. If it wasn't for Gemini/Antigravity I'd have to go with a Max Claude plan, as it stands now I can get by with just a Claude Pro plan to get Opus when I need it, whilst using Antigravity as my day-to-day workhorse. Unfortunately Gemini Flash became too expensive to use as a general purpose model (i.e. for AI features in Apps), luckily there are plenty of cheaper Chinese models to fill that gap now.
- ValentineC 2mo agoWhy do you think it's the best plan available?
- JacobAsmuth 2mo ago(Rate limits * capability of the model) / cost
- Andrex 2mo agoWhat's the current outlook on Antigravity IDE vs. 2.0? How long will they begrudgingly keep it going before kicking everyone to 2.0/3.0? (I actually use a mix of both for some offline projects, nothing serious.)
- mythz 2mo agoAntigravity is now split into 2 Apps: 'Antigravity' which is an agent-first editor layout where you don't see your code and just prompt it. 'Antigravity IDE' which uses the Windsurf/VS Code editor, which is still what I primarily use in my day-to-day. I hope they never retire the IDE, I don't think I can get used to prompting an AI Agent without being able to see my code to help workout what needs to be done.
- GodelNumbering 2mo ago[dead]
- lenerdenator 2mo agoWe're almost five years into the whole GenAI thing and we're still relying on these guys to spoonfeed us incremental updates. It's time for them to start focusing on open-weight models and efficiency. Otherwise there's just a layer of marketing hype and "will it do this?" that has to be cut through for evaluation of each and every release cycle. Models are getting easier and easier to create. The money, if there's any here, is in the harness the user interfaces with, and the data centers running them.
- raffael_de 2mo agois it just me or is this one-upping each other every few days getting ridiculous secreting a whiff of desperation?
- JacobAsmuth 2mo agoJust you. This is typical market competition in a fast moving field.
- vlad_recomply 2mo agoModels are expensive and low performance. On top of that they make you jump through hoops to even use these models without being throttled even for the weaker models. The only reason we are using them is credits. As soon as credits run out we are switching immediately.
- veloxxn 2mo ago[flagged]
- QuesnayJr 2mo agoI remember back when Gemini looked like it was the best model that this comment section was full of confident predictions that Google had "won" and that no one would ever catch up with them again. The most embarassing part is that I kinda believed them.
- sagex 2mo agoDon't know why are they even pursuing Gemini. Just download the Kimi, call it Kimini and serve it on your GPU. Maybe then train next architecture based on this!
- summerlight 2mo agoLooks like 3.6 Flash is the first model with their newest pretraining run (cutoff date is 2026/03), long after 2.5 series.
- XCSme 2mo agotl;dr: 3.6 flash is a bit smarter than 3.5 flash, but also a bit more expensive. My results [0] put Gemini 3.6 Flash at the top. 3.6 Flash high has same $1.5 input price as 3.5 Flash, but output is cheaper from $9.0 to $7.5. Google said 3.6 Flash is more token efficient, but in my tests it's actually LESS token efficient[1] than 3.5 Flash, so despite the output price reduction, it still costs more. [0]: https://aibenchy.com/compare/google-gemini-3-6-flash-medium/google-gemini-3-5-flash-lite-high/google-gemini-3-5-flash-high/google-gemini-3-5-flash-medium/ https://aibenchy.com/compare/google-gemini-3-6-flash-medium/... [1]: https://aibenchy.com/compare/google-gemini-3-6-flash-high/google-gemini-3-5-flash-high/google-gemini-3-6-flash-medium/google-gemini-3-5-flash-medium/ https://aibenchy.com/compare/google-gemini-3-6-flash-high/go...
- thebigspacefuck 2mo agoIMO Gemini has the best free tier models/app for everyday use. Muse-Spark is perhaps just slightly better, but has none of the connectivity to my GApps (for things like “create a recipe in my Google Docs from this image”). Plus they are probably running these things on every Google search so saving tokens is a huge win for them.
- copperx 2mo agoFree? Did I misread the pricing details?
- thebigspacefuck 2mo agoFree plan, the default tier without requiring a subscription. If you use through the Gemini App or gemini.google without paying anything, the model used is 3.6 Flash. Rankings for text are here https://arena.ai/leaderboard/text https://arena.ai/leaderboard/text For comparison of Free Tiers: - Gemini serves 3.6-flash (rank 12) - ChatGPT serves 5.5-Instant (rank 23) - Claude serves Sonnet 5 (rank 27) - Meta AI serves muse-spark-1.1 (rank 5) While Meta AI serves the better ranked model, it doesn't end up working that well for other things. For example, if I ask "help me buy a new raincoat", it ends up suggesting a Cambodian website, whereas Google is well integrated with Google shopping. It doesn't have the same integration with GApps outside of Gmail/Calendar. A few other email connectors are available. Claude has one of the best interfaces with connectors, skills, and plugins galore, but the model and limits are restrictive on the free tier. Gemini, as far as I know, I've never hit a rate limit on Flash. I believe Gemini is going to gain market share through the free tier funnel while serving models as cost-effectively as possible. People are going to use Gemini because they use GApps and Google. ChatGPT and Anthropic are going to be competing for the API/Business users, but for everyone else they are going have to become Google before Google becomes them.
- TheAtomic 2mo agoI have liked using their consumer products but they don't make it easy, that's for sure.
- game_the0ry 2mo agoAt this point, I think google should consider becoming a hyper scaler for anthropic and open ai, and I predict that that is exactly what they do. The model is no longer the most valuable part of the stack.
- dakolli 2mo agoI use 3.5 flash 10x more than any other model, despite have access to all of them. If I'm going to play a slot machine, I'd rather get the pain over with quickly.
- maxdo 2mo agoquite a good model, the speed/price/quality ration is a new golden intersection for me, not sure if its as good as grok 4.5 but quite fast/capable model.
- deleted 2mo ago[deleted]
- luciana1u 2mo ago[flagged]
- CurbStomper 2mo ago[dead]
- JeremyHerrman 2mo agoGemini 2.5 Flash-Lite has been my go to for cheap document processing at scale (especially with 50% off batch mode), but they are really boiling the frog with pricing increases with each version: gemini-2.5-flash-lite: $0.10 input / $0.40 output gemini-3.1-flash-lite: $0.25 input / $1.50 output gemini-3.5-flash-lite: $0.30 input / $2.50 output (a 6.25x increase over 2.5!) Now watch them deprecate Gemini 2.5 Flash-Lite in the coming months...
- tjwebbnorfolk 2mo agogemma4 is the same price as 2.5-flash-lite, and performs better.
- JacobAsmuth 2mo agoHow has your experience been with Gemma 4?
- Havoc 2mo agoFlash Lite: 0.3/m and 2.5/m Deepseek Pro: 0.435/m 0.87/m That's wildly ambitious pricing by Google. You can maybe get away with spicy pricing at the SOTA edge but at the lower tiers everything is a lot more price sensitive.
- JacobAsmuth 2mo agoYou need to compare cost per task buddy boy. Cost per token doesn't tell you much when you don't know how many tokens a model will use to accomplish a task
- Havoc 2mo ago>boy Seriously?
- JacobAsmuth 2mo agobuddy man
- spstoyanov 2mo agoGlad to see the price is going down but it's still too high for a "fast" model
- 5701652400 2mo agoif only DeepSeeek supported vision, would never use Gemini.
- vinhnx 2mo agoFor anyone wanting a faster overview: I ran the Gemini 3.6 Flash and 3.5 series release notes through NotebookLM and generated a short video summary. Link: https://www.youtube.com/watch?v=SUFBhvQ2tY4 https://www.youtube.com/watch?v=SUFBhvQ2tY4
- speak_plainly 2mo agoIt feels like AI is going to be the end of Google. The post-Schmidt company culture cannot produce consistent, consumer-friendly products that any sane person would want to use consistently.
- fuomag9 2mo agono actual cyber model release, useless
- lukewarm707 2mo ago"The model will be exclusively available to governments and trusted partners via CodeMender soon as part of a limited-access pilot program" we are stealing plutocracy from the jaws of emancipation. i don't want to live in a world where abundance is guarded and shared among politicians and cronies, whilst the rest are left to rot.
- ewaewaewa 2mo ago[dead]
- MILP 2mo agoI'm a big fan of the Flash-Lite models. They're exceedingly fast and deliver great outputs for high volume use cases where you need to process requests at scale. Can't wait to try the newer version.
- imagetic 2mo agoIf only I could use Pi.
- kzrdude 2mo agoWell, you can use the google models from Pi. Go to aistudio.google.com and set up an API key. There is a free quota, it's relatively large for the Flash Lite models. ...and last time I looked the limits were more generous for Gemma 4 there, but they have been tightened a bit. That's how it goes, always changing.
- goldenarm 2mo agoLLM reception is truly extreme, even worse than AAA game releases. Ever frontier lab lived it at least once : missing the frontier by a few months triggers extremly negative reactions, then you take back the lead for 2 weeks, and the hype cycle repeats.
- Andrex 2mo agoI spin a mental roulette on whether the reception on a new release will be "OMG best model by far, no one will be able to catch up for months!" or "OMG this is already outdated, RIP company X, they might as well just give up now, there's no coming back from this." It's quite a fun game. I click into the comments and see if the roulette wheel was right.
- ianberdin 2mo agoPelican svg and a near-perfect 3D MacBook at max effort for $0.16, about a fifth of Fable's price. Fable 5 still wins on detail with no visible errors, but it's close. And this isn't a memorized pelican; https://playcode.io/blog/macbook-svg-benchmark#gemini-3-6-flash https://playcode.io/blog/macbook-svg-benchmark#gemini-3-6-fl...
- Arshad-Talpur 2mo agonever tried gemini for coding, but this news seems to be compelling, i would definitely give it a try
- sega_sai 2mo agoI have just tried to switch to 3.6 instead of 3.5 in antigravity and it seems to constantly spit "critical instruction: STOP CALLING TOOLS NOW. YOU MUST WAIT FOR WAKEUP. ". I think I will switch back to 3.5
- parasti 2mo agoKind of excited about this. 3.5 Flash on Antigravity has surprised me recently on a hobby project. When given opportunity to plan, it can deliver on tasks that would take me a while on my own and generates responses at blazing speeds - compared to what I'm used to at work with Opus 4.8 (granted I don't use Opus 4.8 on my hobby projects so just anecdotal). While with Gemini CLI I would just watch it run in circles and run out of 5h allowance before anything useful is produced (or even approached).
- lilytweed 2mo agoReally, what's up with Gemini still not supporting connectors/MCPs/plugins/whatever-they're-called-this-month on web? It makes it a non-starter for any kind of serious use.
- nicce 2mo agoWhy would you use them on the web? Serious use happens elsewhere.
- vrosas 2mo agoI have no skin in this game and this comment will be gray in a few minutes BUT a friendly reminder that these types of threads are astroturfed heavily by competitor labs and any info should be taken with a massive grain of salt.
- arjunvrofficial 2mo ago[dead]
- lwansbrough 2mo agoPlugged 3.5 Flash Lite into an existing agent harness that was previously using 3.1 Flash Lite and this shit just does not work. It's not following instructions and is not producing the correct tool calls.
- kzrdude 2mo agoMaybe this is relevant? Just in case > For autonomous subagents with tool calls, code execution, or multi-step reasoning: set thinking_level to "medium" or "high" to prevent premature tool termination. I just happened to see that in the docs: https://ai.google.dev/gemini-api/docs/latest-model https://ai.google.dev/gemini-api/docs/latest-model
- lwansbrough 2mo agoThanks, I'll give that a try.
- Alifatisk 2mo agoIn other good news "the model has been trained to minimize refusals for beneficial uses.". Otherwise, this news feels like a tiny incremental improvement on Gemini Flash series to make it more efficient with token usage, subagent and cost. Nothing big. Regarding their benchmark scores on CyberGym, I wonder why they didn't compare their 3.5 Flash Cyber model with Fable 5. I mean they included Mythos and GPT-Cyber, so why not Fable 5 too? They also mentioned Gemini 3.5 Pro is in testing and its about to become available very soon. Another thing maybe worth discussing is the announcement of pre-training Gemini 4. Sadly, not much technical details to discuss on. Many comments in here seem to mostly be about how Google is behind the others, but honestly, is it really worth the investment to be #1 in Artifical Analysis every week?
- ur-whale 2mo agoWhy exactly are they announcing these completely milquetoast models ? I'd be low-keying the release if anything, given how lame they are compared to their competition. What am I missing?
- arjie 2mo agoTheir naming scheme is confusing. Branding has never been Google's strong suit and their marketing copy is pretty bottom-of-the-barrel[0]. Anthropic has a pretty clear set of models but Gemini decided to rebrand their Flash as Flash Lite (and presumably the future will see a Flash Lite Mini, a Flash Lite Mini Nano and a Flash Lite Mini Nano 3B) which confuses the pricing to high hell. This plus the Vertex, AI Studio, Gemini, Antigravity. It's honestly too confusing to use. I need to use Gemini just to decide on which platform and which model to consider. 0: Famous Kurian Tweet: "We're announcing Duet AI for Google Workspace will now be Gemini for Google Workspace. Consumers and organizations of all sizes can access Gemini across the Workspace apps they know and love. We're introducing a new offering called Gemini Business, which lets organizations use generative AI in Workspace at a lower price point than Gemini Enterprise, which replaces Duet AI for Workspace Enterprise."
- mrandish 2mo agoI often use Gemini free web chat because it's generally quite good at web search-related questions (apparently it has direct token-level access to the Google Search index) but I noticed in the last two weeks output quality of 3.5 Flash seriously degraded. Maybe they were switching over systems.
- culi 2mo agoGoogle's Knowledge Graph is a massive advantage no other competitor has. I don't think they've fully utilized its full potential but I don't know if any other company could've built something like Scholar Labs
- Andrex 2mo agoKnowledge Graph + automated transcriptions of almost every YouTube video = giant untapped moat of data
- weird-eye-issue 2mo agoIt's not exactly untapped, my AI company has scraped YouTube transcripts for 3 years now for RAG
- culi 2mo agoYeah and I've used this site for years when trying to recall a lecture I watched but only remember bits of https://filmot.com/ https://filmot.com/ (it lets you search youtube by transcripts)
- Andrex 2mo agoThat's true, I meant more by Google themselves. It's their first-party data, the whole set available on-demand, no scraping needed.
- 1saadcodes 2mo agoNice to see that it's cheaper than 3.5
- mchusma 2mo agoWow, Laguna S 2.1 (released today) just destroys Flash-Lite underly and completely. What a weak and embarrasing release from Google.
- marshalla 2mo ago[dead]
- hmokiguess 2mo agoSpent half an hour just now benchmarking it against my current 3.5 Flash pipeline excited only see it regressed slightly (0.1% - 0.2% at most, for feature extraction work) Seems like this is mostly a cost play by Google, hoping this doesn't bring 3.5 Flash capabilities to an end of life, and that 3.6 catches up or gets better.
- waldrews 2mo ago3.5 Flash-Lite seems available in US region, as was 3.5 Flash; but 3.6 Flash looks Global only so far when pinging. If Google employees are watching, will this issue go away?
- jdthedisciple 2mo agoBottom line it looks about on equal footing with GLM 5.2 in terms of both overall intelligence and cost per task, while being significantly faster (in fact it is the fastest model on artificial analysis as of rn [0]) [0] https://artificialanalysis.ai/models/gemini-3-6-flash https://artificialanalysis.ai/models/gemini-3-6-flash
- DekryptLabs 2mo ago[flagged]
- brap 2mo agoFrom my experience, this thing is crazy fast. Spawn 10 on the same problem and have them debate to reach a consensus, you’ll get Fable-like results but 100x faster.
- zwaps 2mo agoHere's the issue: GLM 5.2 is better, also cheaper, and almost as fast. So essentially, a big L for Google. Combine this with them not being able to produce a frontier model this generation... hmm implications
- WarmWash 2mo ago3.6 is roughly 50% faster, which isn't totally insignificant for being marginally more expensive.[1] [1]artificialanalysis.ai
- zwaps 2mo agoSure, but there's no sota alternative from Google. That's it, and its beaten by GLM 5.2 on every measure except somewhat speed. I find that quite staggering. GLM is open weights
- Narkov 2mo agoSpeed is definitely a marketable quality. All these things are a trade-off and solely measuring against SOTA I don't feel is always helpful.
- HDBaseT 2mo agoCounter-point, cost per task is almost the same ($0.47 vs $0.50) between GLM 5.2 and Gemini 3.6 Flash. [0] Not to mention the subscription plans likely produce 10x value compared to GLM 5.2 API, unsure the rate limits on a equal subscription vs subscription, but Google subscriptions offer tons of other value, including 1 year of Free Gemini for Education accounts. [1] https://artificialanalysis.ai/models/gemini-3-6-flash https://artificialanalysis.ai/models/gemini-3-6-flash
- zacksiri 2mo agoGemini 3.5 flash-lite is more expensive than Gemini 3.1 flash-lite. Every upgrade is getting more expensive.
- michaelbuckbee 2mo agoIt's kind of ridiculous how good these are getting. 3.5 Flash lite is pretty comparable to Opus 4.8 (at least for the couple tests I did) while simultaneously being 6x faster and 19x cheaper. https://fy2zp1ri90.evvl.io/ https://fy2zp1ri90.evvl.io/
- s3p 2mo agoNot for me personally. While setting up a custom website, 3.5 Flash introduced tons of bugs that Claude had to fix. The website has about 3,000 lines of code spread across multiple files, and Gemini somehow couldn't do frontend changes without breaking things. Sharing my 2c, but I've stayed on GPT 5.5+ and Claude Sonnet/Opus 4.6+. Anything past that from those two have been bug-free, but Google's latest hasn't been.
- sajithdilshan 2mo agoGoogle really needs to get their product strategy together. The discontinued gemini-cli and introduced antigravity-cli which is a downgrade IMO and the sooner they can partner up with AWS and release the gemini models via Bedrock the easier corporate/business which has strict data protection rules can use their models and make them available for internal engineers. It's a one thing to research and improve the model, but if they ignore the ease of access and multi-availability of their models in different ways they are going to fall behind again.
- resonious 2mo agoSo it's GLM-5.2 performance for almost twice the price. That said, the speed looks really good. I think it's competitive with Fireworks's GLM 5.2 Fast, although Fireworks is still cheaper.
- prtmnth 2mo agoMy hunch is Google is trying to integrate a fast and relatively cheap AI across search and every other surface of their product suite. And for that objective, a model that can move faster while being accurate and cheap enough is more important to them than producing a frontier class heavyweight model.
- schainks 2mo agoThis. Give me cheap tokens that produce accurate information and the deal is done
- verdverm 2mo agoI've found that having good source material (markdown, dependency source, search results for agents) for the models to draw on significantly improves information accuracy. Definitely worth investing in this side of "harness engineering", don't rely on facts burned into weights
- do_anh_tu 2mo agoMan I love Gemini models but these kind of pricing increase is just insanse. I have a little product and I have to keep increasing the price and reduce the limits because of this non-sense, and they did not even let us use the old models in near future, so I forced to update to the new model with basically no to little improvement because I don't even need that much. Google if you can read this, it okay to release new models and change the price for them, but please please don't kill the old ones like gemini-2.5-flash-lite, because that all I ever need for my little apps with only few thousands of users.
- Tuna-Fish 2mo agoDo not base products on models that are not open-weights. Doing it is like building a product on someone else's platform, you are entirely at their mercy, and even when they don't have any reason to hurt you, you are tiny enough that if any policy they want to enact hurts you as a side effect, no-one is going to care. You don't have to self-host the open-weights model, you just need to be able to source it from multiple providers. Using the closed vendor models maybe made sense when open-weight models lagged so far behind, but that time is now gone.
- do_anh_tu 2mo agoI just have no choice. My application was tested and benchmarked against a lot of models, but no model can be in the same league as gemini in term of multi-language support and understanding.
- satvikpendem 2mo agoWhy don't you just switch to cheaper models? I'm sure DeepSeek is probably enough for you and it's way cheaper. If you want, host your own (or have someone else host) open weight models, I use both embedded and cloud Gemma for some things.
- shaism 2mo agoWhat is your use case? Have you considered moving to open source / Chinese models? If gemini-2.5-flash-lite is good enough for your application, you will find even lower cost options with better performance outside of the Google ecosystem.
- lostmsu 2mo ago3.6 Flash has the same performance on artificial analysis benchmarks as 3.5 Flash. So... what... is... the... point?..
- lostmsu 2mo agoNevermind, I just realized it is point six!
- t2ance 2mo agoWaiting for Gemini 3.5 Pro...
- rayzia 2mo ago[dead]
- csunoser 2mo agoThis is the darnest thing - all the testimonies are just jepgs? Did they run out of time to work on the css?
- ernestrc 2mo agoFable or gpt5.6 sol for planning. Gemini 3.6 Flash for executing. Wow, Google is onto something here. I always thought that gemini 3.5-flash was the most underrated model. Let's see how much better is 3.6 flash.
- feiz45607 2mo ago[flagged]