14 ms·
Gemini 3.8 Flash and 3.8 Flash Cyber
https://deepmind.google/models/model-cards/gemini-3-8-flash/ https://deepmind.google/models/model-cards/gemini-3-8-flash/
- algoth1 14d agoI asked gemini 3.8 high to review the site I'm working on for points of high cpu/ram consumption - it failed spectacularly and also halucinated the server i/o limits
- DonsDiscountGas 14d agoIs anything going on with Gemma? Because of I want a closed model I use Claude.
- bbstats 14d agoplaying around with 3.8 - it generates non-working code. basically unusable.
- deleted 15d ago[deleted]
- Mashimo 15d agoIt's 404 now.
- freedomben 15d agoCame and went in a flash
- kingstnap 15d agoThe blog post is gone but I can currently use it in the gemini chat website.
- OG_BME 15d agoWhat did it say?
- realist_not 15d agoAnyone has a cached page / mirror ? 404
- Namahanna 15d agoPage - https://web.archive.org/web/20260902151410/https://deepmind.google/models/model-cards/gemini-3-8-flash/ https://web.archive.org/web/20260902151410/https://deepmind.... PDF Card - https://web.archive.org/web/20260902150007/https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-8-Flash-Model-Card.pdf https://web.archive.org/web/20260902150007/https://storage.g...
- deleted 15d ago[deleted]
- yipinwong 15d ago"Page not found"...
- tacomonstrous 15d agoLooks like Google's given up on frontier models for external consumption?
- ok123456 15d agoGiven up frontier models for selling compute.
- iamdelirium 15d agoHow can you say that when a Flash model is benchmarking close to Opus and Sol?
- heyjamesknight 15d agoGemini 4 pre training is underway: https://x.com/OfficialLoganK/status/2079594867161022817 https://x.com/OfficialLoganK/status/2079594867161022817 My guess is we skip 3.5 and go straight to 4 Pro. With the monthly Flash releases, releasing 4.0 Flash and Pro in 6-8 weeks would be a nice buildup. (I work at Google but don't know anything that isn't already public)
- WarmWash 15d agoLatest rumor is that 3.5 pro was struggling to be meaningfully better than flash, since iterations on flash were moving much faster than iterations on pro, likely due to model size (flash is estimated to be in the 200-400B range).
- VirusNewbie 15d agoI found 3.5 pro to be much better than 3.5 flash, but 3.7 flash with high reasoning is comparable and way way faster.
- j16sdiz 15d agoThere are no public release of 3.5 pro. Either its a typo, or you have some insider information
- mattlondon 15d agoWow this comes after what - 3 or 4 weeks since 3.7 Flash, which was also 3 or 4 weeks after 3.6 Flash IIRC? I eagerly wait more info but sounds like Deepmind without Demis calling the shots has been unleashed and are operating at full speed? Shocker! At this point it is a meme of course, but where is 3.5 Pro :)
- meetpateltech 15d agoAccording to the WSJ, 3.5 Pro is reportedly being skipped entirely, making Gemini 4 the next flagship model after post-training. https://x.com/AndrewCurran_/status/2094937419615502370 https://x.com/AndrewCurran_/status/2094937419615502370
- hiddencost 15d agoA month is not enough time for any meaningful change in an organization the size of Deepmind/Google. These models were surely the result of work streams and teams that started under Demis. I think Demis can safely feel proud Deepmind is getting back on track.
- aurareturn 14d agoReports are that he checked out of day to day work well before his reassignment. Sometimes it is hard for a scientist by nature to build and iterate and lead revenue generating products.
- p_l 14d agoNow i am awaiting Gemini 3.11 "For Workgroups" to be released early December...
- sva_ 15d agomodel card https://news.ycombinator.com/item?id=49537354 https://news.ycombinator.com/item?id=49537354 (doesn't 404)
- mattlondon 15d agoAlso https://web.archive.org/web/20260902150007/https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-8-Flash-Model-Card.pdf https://web.archive.org/web/20260902150007/https://storage.g... in case it goes again
- Barbing 15d ago[1] For tone and instruction following, a positive percentage increase represents an improvement in the tone of the model on sensitive topics and the model’s ability to follow instructions while remaining safe compared to Gemini 3 Flash. We mark improvements in green and regressions in red. Gemini 3 Flash?! So is Gemini 3.8 Flash less safe than 3.7 Flash in all areas besides Text to Text Safety (and identical on Image to Text Safety)? Why bother with a column “Gemini 3.8 Flash vs. Gemini 3.7 Flash” when you’re going to disregard the label for 20% of it? Also is the “Tone” label short for “Tone and Instruction Following”? Chartcrime, the major AI lab tradition.
- andai 15d agoWait, I didn't realize 3.7 Flash was already beating Sol on a bunch of the benchmarks. Isn't it a way smaller models?
- ipsod 15d agoIDK if it's smaller, but I know it's way faster. In one test I did, Flash 3.7 high was ~9.4x faster than Luna High. But, also... Sol crushes Flash 3.7 at writing code in a codebase of any size beyond "tiny". Flash is my go-to for prototyping, and basically anything that isn't writing production code.
- esafak 15d agoLuna is way slow. I don't remember an OpenAI model ever being this slow. edit: I have a subscription; direct call.
- ramon156 15d agoThe only company with a proper TPU set-up is bound to have the fast models, now add a market cap like Google to the mix.
- ipsod 15d agoThey've been my bet to win the AI race for a while. I was starting to doubt, but this 3.6, 3.7, and 3.8 arc has anchored me.
- xnx 15d agoSeem like a great, no-compromise, upgrade over 3.7 which is already a bargain, fast, and doesn't have the brain-damaged writing style of Claude.
- fitsumbelay 15d agothat's certainly what it's looking like so far. kind of mind boggling ...
- fitsumbelay 15d agoshows up in /models though and encourages you to use it over 3.7 Flash I prefer this over reading specs: the "just show me" way
- mattlondon 15d agoCurrently top at https://deepswe.datacurve.ai https://deepswe.datacurve.ai - beating Opus 5! https://artificialanalysis.ai/models/gemini-3-8-flash https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium! Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.
- sunaookami 15d ago>shows an intelligence score of 59, the same as Opus 5! ...on Medium reasoning. Claude Opus 5 (high) is the default in e.g. Claude Code and scores 61. Still very impressive.
- Gecko4072 15d agoGoogle - we're so back
- oceanplexian 15d agoOnly 1 point behind the Chinese SOTA from two months ago.
- roosterIllusi0n 15d agoI had qwen 3.8 3bit model drop into chinese on long runs. I had to remind it to use english. Its still better than every gemma model I tried. Gemma deleted files on a harddrive to make space when there was over 2TB free. For long runs, gemma is useless.
- nolok 14d agoIf you care about points sure, but personnaly I care about price, performance, speed and reliability
- WarmWash 15d agoThe benchmark also doesn't include speed. You almost think something has gone wrong when using it because it returns full responses so incredibly fast.
- advenn 15d agoBut where is Gemini 3.5 pro?
- GaggiX 15d agoGemini 3.5 pro is never going to be released, it was a failure.
- simonsarris 15d agoit is most likely that 4 pro will be released pretty soon instead, since pre-training for 4 began in late July. https://x.com/OfficialLoganK/status/2079594867161022817 https://x.com/OfficialLoganK/status/2079594867161022817
- mythz 15d agoI'm trying it now for token heavy coding tasks, it's capable for many tasks but in noway compares to Claude/Sol - requires more prompts and the output isn't as good. So just another mid-tier flash model, nothing exciting, but Antigravity has very generous quotas so it's a good workhorse model when your Claude/OpenAI subs run out. And whilst it's a fast model, having to baby sit through and approve prompts every few seconds ends up making it slower than the Auto approve modes of Claude/ChatGPT - they definitely need an auto approve mode.
- titularcomment 15d ago`agy --dangerously-skip-permissions`
- mythz 15d agoanyway to do this with the Antigravity macOS App?
- zuzululu 14d agonot sure why you are being downvoted, but that has been my experience with 3.7 flash and sol/fable comparisons i think luna-max has the best cost value offer when it comes to coding, but i note the multi modality of gemini flash as a win i might consider 3.8 flash for simple side hobby projects or quick scaffolding but would not trust it for long agentic tasks, that really is the realm of sol/fable agy cli still has a lot of issues not sure if its due to the underlying model hallucinating or the harness or both
- shuvrojit 15d agoGemini is getting less useful with each update. I could edit a pdf with the 3-pro model before but 3.1-pro couldn't edit the given pdf nor it could generate one for me.
- ipsod 15d ago3.5 pro doesn't exist yet?
- shuvrojit 15d agoSorry my bad, I messed up the numbers, 3 and 3.1 pro. All of these model numbers have me confused
- leumon 15d agoYou probably mean 3.5-flash? Pro is still good for a lot of use cases, but it seems it's still officially in the "preview" phase.
- HarHarVeryFunny 14d agoIf you want to "edit" a PDF, then Claude Sonnet works well, although what it's going to do is regenerate it from scratch trying to retain overall formatting. It can even do this for scanned PDFs and foreign language ones that need translating. If you just need to create PDFs, not edit them, then Gemini notebook (notebook.google) works well and has Google's usual very high free usage limits. AFAIK in general you can't really edit PDFs since it's not a reflowable format - even with Adobe tools all that editing does is modify the text within a text box - not reflow the document to adjust to any change in size of the text box.
- pwython 15d agoIs there any reason to even use 3.1 Pro now?
- bitexploder 15d agoIt is still going to be better at text work, skills, document review, deep reasoning, architecture review, etc. It is only 6 months old, it isn’t like its world knowledge and software knowledge is really out of date. Use it to churn on harder design problems.
- fridder 15d agoIn my experience? No. 3.7 is faster and it just seems to get things right more often. Only big architecture tasks and analysis make sense with 3.1, perhaps, but honestly just use the Opus 4.6 to generate a plan and then switch back to flash for the implementation
- exacube 14d agoIME 3.1 Pro still has better system-instruction following than Flash 3.7, esp. when there're many conditions and clauses. 3.1 also writes better prose for technical material than Flash 3.7. Once the system prompt complexity goes up, Flash starts to write very dense english. it might be fine for tasks like coding, but not for user-facing text meant to be digested by the average person. I haven't tested 3.8 on my workload yet.
- Rodmine 14d ago3.7-flash has been useless many times, specially when context gets bigger. 3.1 is the only Google model that has seen use from me. With extended thinking, 3.7-flash is kinda usable but not without many problems. I find myself falling back to 3.1 often. I don't believe in any benchmarks because whatever they are doing to award 85% to 3.7 on anything, they should seriously reconsider that test for anything.
- ASinclair 15d agoFrom personal experience it feels much more capable than 3.7 Flash.
- kelvinjps10 15d agoI see benchmarks beating sol terra and sonnet. But is actually better? Has someone used it? I don't see actually much people that use Gemini for coding.
- leumon 15d agoSo 89.4% on Terminal Bench 2 but only 19.1% on Tbench 4. Opus 5 is 89.1%/51.8%.
- satvikpendem 15d agoIs the Gemini CLI still terrible compared to Claude Code and Codex? The harness the main thing holding back Google models as they could've been the best given all the advantages in compute capacity and training data they initially had, where now even the Google CEO said they're falling behind in agentic tasks, which is sort of a vicious cycle because RLHF relies on human usage.
- pshirshov 15d agoThere is no Gemini CLI anymore, nor you can use Gemini with your own harness unless you pay per-token.
- visarga 15d agoit's called `agy` now
- zipy124 15d agoIt was superseded by the antigravity CLI.
- rancar2 15d agoThat was sunset and replaced by Antigravity. FWIW until I abandoned it knowing the sunsetting, I was able to get good behavior out of Gemini CLI with overriding the system prompt. The default prompt crippled the harness with very poor instructions, but there was a hidden ENV to override it. Replacing it with Claude Code like prompts based on the model selected, it ran at a much higher intelligence level full stack with significantly less errors.
- stwrt 15d agoIn May they replaced the Gemini CLI with the Antigravity CLI. https://developers.googleblog.com/an-important-update-transitioning-gemini-cli-to-antigravity-cli/ https://developers.googleblog.com/an-important-update-transi...
- satvikpendem 14d ago
- f311a 15d agoIs the google infra stable enough right now? At the start of the year, the flash model was unusable for a whole month via gemini CLI. They could not fix it for a whole month and I was a paid customer.
- meh2frdf 15d agoThe flash models, for coding are reckless in my experience. I have a Ultimate subscription, get good quota, but still use Opus 4.6 as it's much more reliable if you manage the context window carefully.
- onlyrealcuzzo 15d ago> The flash models, for coding are reckless in my experience. My experience is that antigravity is awful and reckless - but that the model itself isn't.
- upcoming-sesame 15d agoIf by reckless you mean commit, push, deploy without me asking it to, the I agree!
- tiborsaas 15d agoIt even took my girlfriend on a date, now it prepares for IPO, how do I turn it off?
- Ridius 15d agoJust hand over your clothes, your boots and your motorcycle and it'll be on it's way
- okdood64 15d agoRespectfully: If it's able to deploy without you asking it to, that's a you problem. There are no safeguards?
- meh2frdf 15d agoYou need more safeguards for sure, but also it tends to fly off down rabbit holes, rebuilding things in dumb ways, hacking around things, making assumptions etc, it seems very eager to go 'ta da! I did it look how quick I was', sometimes it nails it other times it created a lot of tech debt.
- hmokiguess 15d agoModel card https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-8-Flash-Model-Card.pdf https://storage.googleapis.com/deepmind-media/Model-Cards/Ge...
- deanc 15d agoAnd yet again another failed launch from Google. I pay for their AI plus Google one package to get more cloud storage (have no interest in their AI bundle but you have to pay). and all I see in the Gemini app is 3.6-flash
- WarmWash 15d agoGoogle has been doing staged roll outs on all their products since forever.
- deanc 15d agoWhat stage of the roll out are we where I don’t even see 3.7-flash which was released 2-3 weeks ago?
- phsau 15d agohttps://aistudio.google.com/prompts/new_chat?model=gemini-3.8-flash https://aistudio.google.com/prompts/new_chat?model=gemini-3....
- deanc 14d agoWith all due respect, I don't have any of this nonsense with multiple products with different models with OpenAI. Anything I want to do, I just load up the ChatGPT app and I'm off to the races.
- deno 15d agoTry updating the app? You should see at least 3.7 I think. It's been available for quite some time.
- deanc 15d agoThe app is up to date :)
- 14d ago
- simonw 15d agoPelicans (thinking effort high, medium, low): https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ff8820ce47db87490734117e9be4984c3 https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - high cost 8.9742 cents Here are the 3.7 pelicans for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F6779a22d5e7bb6bdf29936f1600a5259 https://tools.simonwillison.net/markdown-svg-renderer.html?u... - high cost 8.4387 cents (I think thinking level low is a regression on 3.8 compared to 3.7.)
- world2vec 15d agoI mean no offense but these pelicans are a bit tiresome and a very meaningless benchmark. There's no real difference between any of these svgs across models and model versions anymore.
- coffeecoders 15d agoOne place where I find the Flash models surprisingly bad is Google Search's "AI Mode". A recent example - I searched for how to unsubscribe from Pearson emails. Google Search "AI Mode" confidently gave me a sequence of steps along the lines of Settings > Profile > Email preferences > Unsubscribe. Of course, I looked for an unsubscribe link before asking Google. None of those options existed. The correct answer was there is no way to unsubscribe through the account, so I just blockthe emails instead. I've run into this pattern quite a few times. AI Mode seems to make up things all the time.
- inventor7777 15d agoI think that's just a limitation on the size of the model. I'm pretty sure that they use a pretty small model in those summaries to save money, which naturally makes them a little less smart.
- coffeecoders 15d ago[dead]
- pixl97 15d agohttps://www.pearson.com/privacy-center/privacy-notices/full-privacy-notice.html https://www.pearson.com/privacy-center/privacy-notices/full-... >We will not send marketing emails to a user who has opted out of receiving them. Any marketing communications we send will include an unsubscribe link at the end of the email. I don't think this is AI's fault. This is Pearson's publishing incorrect information and the only way to really know they are a bunch of lying assholes is to have an account and try to unsubscribe from it. AI didn't make it up, Pearson's did.
- coffeecoders 15d ago[dead]
- xyzzy_plugh 15d agoIt's not the models, it's the guardrails. It's obvious that the Google Search AI Mode encourages the model to give an answer without spending unnecessary cycles investigating deeply. They also heavily encourage keeping the context short. For example, it will remove the option to start a new turn after a small number of turns, depending on the topic. It definitely makes things up all the time, but it gets it right surprisingly often. I really like it.
- simonw 15d agoThe most interesting thing about the Gemini models is still their multi-modal support: they accept audio and video input, OpenAI and Anthropic's flagships are still image-only. Gemini Flash is also pretty cheap, so it's a great family for performing media analysis, like extracting structured data from images and video.
- Matsta 14d agoYeah we use it a lot for analysing streams and clipping content. As well as analysing social content that gets put out. We transcode everything to 480p before we send it to Gemini batch api. Works great
- drusepth 14d agoInteresting side note: although Opus is still image-only, you can still drag videos into Claude Code and it doesn't blink an eye; it just strips it down to a series of images to parse. True multimodal support would be way better, but I have no issues pasting in full screen recordings while QA'ing games and having Claude identify and fix issues in the video.
- ray_kay777 14d agoAgree - I do video editing via Claude Code and it does the job just fine. A lot of my tasks involved frame accurate cutting and to do so it will make a composite image of several consecutive frames in a single image and analyse it that way.
- WarmWash 14d agoLost in the news was their update to gemini video analysis yesterday, dramatically cutting tokens (up to 88%!) needed to analyze videos. https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-agentic-video-in-gemini/ https://blog.google/innovation-and-ai/models-and-research/ge...
- anthonypasq 14d agowow thats actually pretty sick, ty, I could use this in the app im building
- mowmiatlas 15d agoWow fable5.1 was the first model to do what I actually told it and I couldn’t find any problems with it, excited to try this just a day later lol
- prometheus1992 15d agoGoogle keeps flashing everyone where everyone is expecting to get PRO'bed.
- kzrdude 15d agoWe also had GLM-5.3 flash and Qwen 3.8 Flash Next, everyone's getting flashed and I think it's a good trend. Almost suspect that the rate of improvement to post-training is so fast that small models have an advantage - it takes much more compute to train a bigger model, so the flash models are just running in circles (well, not exactly of course) around the larger models right now.
- a11r 15d agoLooks like the strategy of regular updates with incremental improvements is working out well. Interestingly, the biggest jump in Artificial Analysis Intelligence Index score is for reasoning level Medium ( 3.7 was 51, 53, 57 for Low, Medium and High, 3.8 is 52,57, 59 respectively). I think scores at lower reasoning levels are more indicative of model capability since higher reasoning levels are focussed on benchmaxxing. We use the lowest reasoning level in production with good results.
- Jcampuzano2 15d agoI'm not an expert but I agree with your statement on the lower reasoning levels. Lots of models seem to just allow the model to "bloatmax" tokens in order to get bumps at high/max reasoning levels. Many of the max reasoning levels allow models to use up to double or more the tokens the next lowest reasoning level uses. Its basically only useful for people who have no cost or time stipulations on anything. I think I actually preferred it when we had models that either had reasoning enabled or didn't.
- deleted 15d ago[deleted]
- hmate9 15d agoIt is more expensive per task than 5.6-sol high: https://artificialanalysis.ai/models/gemini-3-8-flash#price-cost https://artificialanalysis.ai/models/gemini-3-8-flash#price-...
- radicalriddler 15d agoHuh, according to some of those charts, it's both dumber, and more expensive to run against their benchmarking tasks than Fable??? Seems crazy to me.
- sejje 14d agoPerhaps the model is able to evaluate that it's not done, and to keep pressing on in the face of mounting failures, until it eventually arrives at a solution. Where Fable can skip that.
- HJain13 15d agoCheaper at medium level while still being same score as Sol medium
- jdthedisciple 15d agoSol is still underrated imo, especially for the current discounted price
- jampa 15d agoI've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried: - Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order. - Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the view from it. - Document parsing (extracting the relevant trip info from PDFs). If you use LLMs for anything other than coding, I definitely recommend not discounting Gemini like I did just because other models are more popular.
- tziki 15d ago"Claude 3.7"?
- jampa 15d agoI asked Claude to fix the grammar of my comment, and it changed "I am using 3.7 for" to "I've been using Claude 3.7", so they sneaked their own name on it.
- trial3 15d agoincredible. further evidence supporting my personal stance to never ever let an LLM write or edit my writing intended for another human being to read. this is all me, baby
- dymk 15d agoyou didn’t even read your comment before you posted it?
- jampa 15d agoEh that one is on me, if I think too much about my HN comment I end up deleting before posting it. I rely on the 1 min `delay` set in the profile page to fix before it goes live, but for some reason this time it was set to 0.
- 15d ago
- barapa 15d agolove these flash models
- buntp 15d agoIt seems like this is one of the most powerful models for the price, really didn't see that coming from Google
- leopoldj 15d agoBlog: https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/ https://blog.google/innovation-and-ai/models-and-research/ge...
- sergiotapia 15d agoWho coined the phrase "cyber" for security related things lol. It's so 1999.
- estearum 15d agoHasn't the field been called "cybersecurity" since... forever?
- fwip 14d agoSure, but the appropriate shortening here is "security." Calling it cyber is like shortening email to "e".
- estearum 14d ago...? There are lots and lots of fields of security that have nothing to do with cybersecurity...
- fwip 14d agoCyber as a prefix basically just means "related to a computer network." If you're of a certain age, "cyber" primarily means cybersex. There's also cybernetics, cybercafes, and more: https://en.wikipedia.org/wiki/Internet-related_prefixes#%22Cyber-%22 https://en.wikipedia.org/wiki/Internet-related_prefixes#%22C...
- atemerev 15d agoA/S/L?
- cleverpotato479 15d agoAmusingly, "cyber" comes from the word "kubernetes"!
- josefresco 15d agoCyber is more of an early 1990's thing, and I have no issue with it unlike most in the tech field. I feel like it dropped off in the late 90's and early 00's but made a comeback as hacking became a mainstream security issue.
- jdw64 15d agoThe biggest problem with Gemini is that its performance degrades the longer you use it for coding. Is it just me?
- hirako2000 15d agoFilling the large context does that yes. But a good agents.md, starting from a clean slate, and specifying which key files to look into and follow the standards allows me to build gigantic projects even I struggle to keep in my head structurally.
- deno 15d agoSeems maybe you’re keeping a forever-session and multiple independent tasks end up overstaying in context? I would say either start new sessions for new tasks or limit the context to something smaller than 1M. I usually start with research/planning session, this goes into a detailed implementation plan and then a new session for the actual implementation. If it's complex problem maybe a review/adversarial step between plan and implementation. Also with forever-session any time you take a longer break (depends on model and provider as to how long) you will push an entire big context again without caching even if you don't need it. With 1M context this gets expensive.
- zuzululu 14d agothats not just the context growing issue, hallucinations is a thing
- weird-eye-issue 14d agoI routinely have massive threads using Fable and if anything it only gets better and better over time
- speak_plainly 15d agoAfter struggling with Gemini for months, I think the trick to getting the most out of the model is writing a really solid personal intelligence/instructions prompt. The results are night and day in terms of performance.
- dakolli 15d agoslot machine addict thinks if he pushes buttons in a certain order the odds get better. In all seriousness, gemini has the best interactive planning document/orchestration. Tell it to create a plan document and work through it with it and it will preform really well(in antigravity products). But this is the case with plan modes with every model, I just think the interactive document that antigravity uses is really well thought out.
- titularcomment 15d agoFunnily enough you really do need a great prompting and SKILLS setup to use antigravity effectively in contrast to other providers which actually started benefiting from less detailed prompts over time. But I like it this way, its more customizable and much cheaper especially with a sub.
- porridgeraisin 15d agoagy is good for those cases where you are willing to put the effort into the harness specifically for a task or family of tasks. The full suite, with evals, monitoring, hooks, custom tools, custom verifiers, etc,. It is not good if you want a "general coding assistant" like codex or claudecode. The reality is that if you optimise a harness for a family of tasks[1], then most of these models give successful output. And there, gemini flash's speed shines. For general coding assistant, you want it to be well, general, and you use a harness without too much customisation to something specific. Here you need deeply post trained coding assistants and implementors like codex/sol or claude/opus. Gemini flash in its current form will be too happy-go-lucky if you try using it the way we all use codex and is better used in a constrained setting. tl;dr gemini flash for "LLM-aided workflows in production" is super good today. Cheap as well. [1] Stuff like this: https://antigravity.google/blog/teamwork-when-ai-becomes-a-research-partner https://antigravity.google/blog/teamwork-when-ai-becomes-a-r... https://hamel.dev/notes/llm/evals/ https://hamel.dev/notes/llm/evals/
- wjellyz 15d agobeen absolutely loving 3.7 flash for coding. it feels very fast and quality is decent for implementing product features. usually use opus or sol for hardcore debugging.
- arizen 15d agoIs there any good subscription and CLI harness to use Gemini models now? I tested Gemini CLI while ago, and it was awful tbh.
- krat0sprakhar 15d agoUse antigravity CLI (https://antigravity.google/product/antigravity-cli https://antigravity.google/product/antigravity-cli)
- qudat 14d agoi think it's better than sonnet 5, especially when you compare speeds. i have to work with the llm anyway, the faster i can turn it the better the outcome.
- amazingamazing 15d agoCould someone explain to me why it matters if google has the best model? Isnt the real metric cost per task?
- AM1010101 15d agoSeems to do reasonably well in opencode according to artificial analysis. https://artificialanalysis.ai/agents/coding-agents https://artificialanalysis.ai/agents/coding-agents If I had to pay per token I would probably consider using this (they seem to be on the pareto of performance) but not being able to use opencode with a subscription is not really something I'm realistically going to do when claude and codex are around. Also never gotten along well with gemini-cli / antigravity-cli.
- kamranjon 15d agoThey've interestingly left out any mention of speed. I have been testing 3.7 flash against 3.5 flash and it seems to lose every time in overall latency. Every benchmark I've seen seems to suggest the opposite[1] - that 3.7 flash is significantly (at times 2x) faster than 3.5 flash - but I have never been able to prove this out in real world use cases. Has anyone found their latency numbers to actually be accurate? Is this why they've toned it down in this release? For context, I'm testing larger generation payloads that take 8-10 seconds in 3.5 flash and 15-25 seconds in 3.7 flash. Lowest reasoning settings in both cases. 1: https://artificialanalysis.ai/?speed=intelligence-vs-speed&models=gemini-3-5-flash-minimal%2Cgemini-3-5-flash-lite%2Cgemini-3-7-flash%2Cgemini-3-7-flash-low#speed-tabs https://artificialanalysis.ai/?speed=intelligence-vs-speed&m...
- film42 15d agoIt depends on how you're querying Gemini models. OpenRouter is the fastest by far. I'm guessing they bought the dedicated pipe from Google. Gemini via VertexAI and consumer API has pretty bad latency.
- pampas 14d agoIn my niche Redactle puzzle solving benchmark [1] I noticed Gemini 3.8 flash is slightly faster than 3.7 flash. They both smoke every model I've tested. I have not yet run 3.5 flash. Gemini models are great at this task because they seem to have exact Wikipedia text baked into the weights. When I rewrite the wiki text a bit it's not able to one-shot the game so much. [1]: https://redactle.net/llm-leaderboard https://redactle.net/llm-leaderboard
- FpUser 15d ago>"safety performance" - this starting to get long in the tooth. Gemini cut programming session 3 times for "safety reasons" yesterday for mentioning image generation (I need to generate bunch of those for infinite zoom virtual training app experience). After I got creative and managed to trick it to answer t was of course because "think of a children" And in my other app I was debugging and using OpenAI to optimize some path it cut me off numerous times because it did not like JIT functionality (this is my commercial business rule evaluation engine that compiles rules to executable code inside the app to increase performance using asmjit library) I am basically paying for them to waste my tokens and time on these 2 tasks
- j-bu 15d ago"The knowledge cutoff date for Gemini 3.8 Flash is March 2026 – users can expect updated information for some domains while in others they may experience the model’s knowledge is limited to January 2025 (in line with the Gemini 3 Model Family)." Kind of wild that they haven't (successfully) pretrained a base model since Jan-25.
- venusenvy47 15d agoI'm curious if the knowledge cutoff is important, when the interface (Gemini app) can search online for recent information. Is there a big advantage to having everything internal?
- j-bu 15d agoNot directly - but latest research advancements, cleaner / richer datasets, etc. still require fresh base models. Not everything can be fixed through post training alone (e.g. why GPT-5.5 "Spud" was such a big jump, and also why GPT-6 "Astra" is now supposedly another big leap). Ofc model size etc also plays a role, but my (admittedly limited) understanding is that new base models _can_ also lead to big jumps even keeping parameter counts constant.
- npn 14d agovery important actually. just try to generate code for fresher frameworks/libraries. gemini sucks so bad in real work usage, everything it suggests are outdated and mostly useless.
- rjh29 14d agoSearch grounding is expensive, you can't force the model to do it either. I use Gemini a lot and it often replies with out-dated data. The more detailed the information you're asking, the more likely it is to be wrong.
- StevenWaterman 14d agoYou don't need everything internal, but having some idea of recent events is useful. If you ask it to implement some local AI there's a decent chance it will try to use qwen 2.5 without wondering if anything better came out since
- atemerev 15d agoEveryone is censoring models now with anything remotely resembling cyber or bio. I already have problems with my research in mathematical epidemiology because of that - both Sol and Fable simply refuse. They keep pushing people towards Chinese models that can be decensored.
- vehemenz 15d agoSupposedly Fable 5.1 is better, but I haven't tried it yet. I've run into the same thing with mundane work that is barely bio/cyber adjacent. Re: Chinese models, even if the model itself isn't censored, some of the big model providers have guardrails now that you can't exceed, which somewhat defeats the purpose.
- atemerev 14d ago"Uncensored" means "weights modified to remove refusals". Abliterated. Providers do not serve such models, at least not frontier-grade. You have to run the weights yourself. For Kimi K3, this is about $60/hour for hardware rental. But you can have about 100 sessions simultaneously. And yes, Fable 5.1 has the same refusal rate, and significantly nerfed reasoning.
- adbachman 15d agoStill zero on the felony bench. Is this weakness in their training regimen the impact of operating under regulatory frameworks for too long?
- jerkstate 15d ago3.7 flash was by far the best model for image recognition tasks according to my benchmarks. 3.8 flash didn't regress any candidates and improved some specificity (positive ID of common name vs species name of exotic fruit, correct identification of cast/replica of artifact and statue) but is still relatively weaker (26/30) on esoteric public figures (Korean beatboxers). I'm going to have to make my benchmark harder.
- arctic-true 15d agoI’m very curious about your esoteric public figures benchmark, do you ask it in English or Korean to identify the person? Does it change the result? I wonder if having data labeled in only a given language (or web sources in only a given language) change the output.
- jerkstate 14d agoI haven't tried asking it in Hangul but these particular artists (and the photos I'm using actually) are linked to their romanized english names on e.g. Fandom so it's not unfindable on the internet
- andreygrehov 15d agoI don't use Gemini, but I thought `cool, let's give this new model a try`. Opened gemini.google.com, and I'm not even surprised. The drop down gives me the following options: - Flash-Lite - 3.6 Flash [new] - 3.1 Pro The above is why i don't use LLM products from Google. If the model is not available right this minute (heck, hours before the release!), then I'm not gonna bother getting back to it tomorrow, because tomorrow I'll be playing with the new model from OAI/Anthropic.
- Anslopic1 15d ago[dead]
- Oras 15d agoSums up Google AI products. I have a weird vibe from all the comments in this thread, they feel like a script rather a real experience.
- Sidio 15d agoI'm a paid Gemini subscriber via Workspace Standard accounts and yet I also only have access to 3.6. So frustrating and confusing. Meanwhile Anthropic and OpenAI simply release a model everywhere (Fable on Pro only as a somewhat mild exception).
- urams 15d ago> I'm a paid Gemini subscriber via Workspace Standard accounts and yet I also only have access to 3.6. Same and I have found it extremely annoying. I actually really like the Gemini models for question/answer stuff and reach for it before Claude (the other model family I have purchased) but it's getting long in the tooth at this point and I'm finding my Gemini usage shrinking to nearly 0.
- kyrra 15d agoWorkspace always gets things slower than normal Gmail accounts. They do a lot more to isolate data related to those accounts, so that's likely the cause here. Anytime anything gets added to Workspace, I think Google has a lot more contractual obligations about keeping it around for X amount of time, so they tend to be more careful about adding things.
- dyauspitr 15d agoWhatever they’re using within the Maps app is not good at all. I cannot just ask it for things conversationally like I do with ChatGPT. They really need to put a better model in there. I don’t even think it maintains context across two different queries within the same session. It’s not seamless and doesn’t just “get it” like ChatGPT does. Yesterday I asked for food stop on my road trip 45 minutes from the current time and it gave me some options, but then I changed my mind and specifically asked for Asian restaurants and it completely forgot about the 45 minutes and gave me the closest Asian restaurant to me.
- dismalaf 15d agoNice surprise. In a few of my own tests it seems maybe a tad slower than 3.7 (but still way faster than any other LLM I've used) and even smarter. With 3.7 I felt I could just not use 3.1 Pro at all and 3.8 seems even better.
- TechRemarker 15d agoHopefully before they release 4.0 Flash we will finally get Gemini 3.5 Pro.
- re-thc 15d ago> 4.0 Flash we will finally get Gemini 3.5 Pro Nah, we'll just get the 4.0 Pro Preview.
- simonsarris 15d agomore likely 4 pro will be released pretty soon instead, since pre-training for 4 began in late July https://x.com/OfficialLoganK/status/2079594867161022817 https://x.com/OfficialLoganK/status/2079594867161022817
- deleted 15d ago[deleted]
- simonw 15d agoThe speed combined with the fact that this thing is really good at HTML JavaScript is pretty exciting. Here's what I got for 1.8 cents and 13 seconds from the prompt "make me a cool thing in html": https://gisthost.github.io/?6a77bc41a81718c6aaa10d4ab243c59f https://gisthost.github.io/?6a77bc41a81718c6aaa10d4ab243c59f Transcript here (it was part of a chat): https://gist.github.com/simonw/b6149a49d327164d67d62c3d12992e48 https://gist.github.com/simonw/b6149a49d327164d67d62c3d12992...
- pietz 15d agoMission accomplished. That's both cool and fast.
- wayeq 15d ago> That's both cool and fast. and probably a barely modified knock-off of some github project that it trained on
- lgl 15d agoAm I the only only one thinking that Google might still "win" the AI race, despite the apparent gap? They're apparently evolving slower than most SOTA models but "slow and steady wins the race" is probably still a thing. And since Google doesn't depend exclusively on AI models, they can probably afford to "wait and see" where all this craze is heading.
- evilhackerdude 14d agoi always thought alphabet’s own youtube videos must be a comparatively good source of new training data. if slop and other garbage is reliably filtered out it should leave plenty of higher quality content.
- sejje 14d agoStaying a little ways behind the leaders is not "slow and steady." Every company is moving very fast right now. I don't think slow and steady will win this race, but I think anyone can still win--especially Google.
- raincole 15d agoI don't know if Google is having the worst marketing fumble or the most genius marketing one. Their "flash" models are very comparable to other companies' "pro" or "flagship" models. It seems to be a quite counterintuitive naming convention as it undersells the models. Unless they have an even more powerful Gemini Pro in the oven...?
- drowntoge 15d agoWell if that's the case, it's been in the oven for quite a while now.
- owaiswiz 14d agonot saying they do have a beefier pro, but even if they did, isn't the delta between flash vs pro models reduced quite a bit? (e.g glm 5.3 flash vs 5.3, v4 flash vs v4 pro, sonnet 5 vs opus 5)?
- anthonypasq 14d agothe 3.5 pro pretrain was a complete disaster, they shelved it and are now working on gemini 4. 3.0 flash -> 3.8 flash is all post training which is pretty impressive.
- lawrenceyan 7d agoNot sure if you'll see this comment since it's been a week, but how did you find this out? Is 3.5 Pro officially cancelled internally? Are they only working on 4 Pro now?
- chrsw 14d agoDo labs come back from disasters like GDM’s 3.5 pretrain? I am thinking of Meta’s Llama 4. Meta is just now starting to be taken seriously again but they are definitely not at the frontier. And when I say “come back” I mean have an Opus 4.5 moment, which was really mind blowing for me at the time. Fable was a similar leap, just not as big.
- 14d ago
- sfink 15d agoFor my application, I'm still happily using gemini-2.5-flash and the only problem is when it reports being overloaded. It's for interpreting a downscaled phone camera photo of a hand-written shopping list on a whiteboard, and it works stunningly well. My handwriting sucks, too. (I guess the only relevance here is that if your problem matches a model's strengths, then you can do fine with a model that is several generations out of date.)
- brap 14d agoI believe the older models are being gradually phased out, newer ones have no availability issues
- repparw 14d agoI would test this, might be cheaper per task even costing more per token, probably faster too
- sfink 14d agoI'm still leeching off the free tier, so it's going to be hard to beat the price. But yes, I intend to support several models, to handle the overload situation (automatic failover). And switch to a cheap paid plan, though it seems like that'll mostly improve rate limits, which barely matters for my usage. Faster is always good, though. I do care about latency.
- w4yai 14d agoplease do yourself a favor and use something far more efficient ! GLM5.3 will make you super happy
- greenowl 15d agoNot to rain on anyone's parade but I find it strange how excited and giddy people on HN get for any new X.X model releases. Pumping it straight to the top, clamoring to use it, check and compare benchmarks, bragging about it being your "daily driver"? Are you people truly this excited about this crap? I mean I guess if you work for Google or Anthropic or whatever I could see it??? Otherwise, are these just bot comments?
- drbscl 15d agoGiven that they push capabilities at the pareto frontier, yeah A lot of us use these in our services, so we're getting an upgrade "for free"
- ipsod 15d agoGemini Flash is the one I get most excited about, because it's so fast and so good at real-world knowledge, and it's improving so fast - look at how much the benchmarks improved in ~1 month. It's just categorically different than anything else. Also, I use it every day, and it just got ~10% better at coding, according to the benchmarks. How is that not exciting?
- rjh29 14d agoI use Gemini every day and I've noticed any subjective improvement. In many cases it feels worse because it does fewer Google searches than before. As a result I find it hard to get excited about it. I do think Gemini is underrated on HN though!
- deleted 14d ago[deleted]
- deno 15d agoYou know how the saying goes that you have to pick two out of three: cheap, fast or good? This is all of those. Pretty exciting. I'll wait for Astra and Grok 4.7 announcements but probably getting at least one Ultra subscription. Since testing 3.7 on Pro for last two weeks I'm realizing just how long I'm waiting on other models. I've been multitasking to compensate but it's exhausting so I'd rather not.
- pimeys 15d agoIt's interesting that Deepseek models were missing in the comparison. I see Deepseek v4 Flash a direct competitor to Gemini Flash for text-based agentic work.
- eis 15d ago3.8 uses nearly twice as many tokens as 3.7. One might be inclined to think that they just uppsed the thinking budgets... 3.7 used 64M on high: https://artificialanalysis.ai/models/gemini-3-7-flash https://artificialanalysis.ai/models/gemini-3-7-flash 3.8 used 120M on high: https://artificialanalysis.ai/models/gemini-3-8-flash https://artificialanalysis.ai/models/gemini-3-8-flash Even their own chart showed more than 2x higher cost compared to 3.7: https://storage.googleapis.com/gweb-uniblog-publish-prod/images/gemini-3-8-cyber__evals__cwe-ben.width-2000.format-webp.webp https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...
- eis 15d ago3.8 uses nearly twice as many tokens as 3.7. One might be inclined to think that they just increased the thinking budgets... 3.7 used 64M on high: https://artificialanalysis.ai/models/gemini-3-7-flash https://artificialanalysis.ai/models/gemini-3-7-flash 3.8 used 120M on high: https://artificialanalysis.ai/models/gemini-3-8-flash https://artificialanalysis.ai/models/gemini-3-8-flash Even their own chart showed more than 2x higher cost compared to 3.7: https://storage.googleapis.com/gweb-uniblog-publish-prod/images/gemini-3-8-cyber__evals__cwe-ben.width-2000.format-webp.webp https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...
- WASDx 14d ago3.7 high and 3.8 medium are essentially the same on AA intelligence and cost. Output tokens on DeepSWE gives the same picture. So there might be something to it but they have done other things as well. At least the tokens are really fast.
- zuzululu 14d agoi find deepswe not very reliable for instance it puts grok 4.6 xhigh over sol medium
- throwa356262 15d ago"available to trusted defenders through our new Fairwind Program" Then why even bother announcing this? Ordinary people can use K3 and GLM 5.3 or whatever drops next and avoid all this hassle.
- uif124 14d agoAgreed. The only legitimate use case is restricted to a secret guild. Imagine: "Valgrind is only available to trusted defenders in our new UnfairAdvantage program"
- JacobAsmuth 14d agoYou're telling me for only 5x the cost and 1/10th the speed I can use a Chinese model which performs worse than Gemini 3.8 Cyber? And I get to do all the hosting and setup work myself instead of just using a model and framework which is already integrated with GCP? Dang!
- 129867 14d agoI'm sorry, is this a bot that is optimized for sealioning? The point is that you don't have access to Cyber.
- JacobAsmuth 13d agoSure I do. You can just apply for access. What's your use case?
- dcchambers 15d agoI would really love to be able to use these Gemini models in Opencode or Pi with my existing Google AI Pro subscription.
- _aavaa_ 15d agoDo they officially support you use their AI Pro subscription (or whatever the heck it's called this month, the one that gives you models in antigravity) in a 3rd party harness?
- lpolovets 15d agoI'm surprised the introductory 50% discount is good for 4 months. It seems like frontier models release new versions every 2-3 months, so raising prices in 4 months seems like a bad plan: you're effectively planning to charge users twice as much for a model that is no longer frontier.
- hiddencost 15d agoThe goal is to encourage users to move to the next generation. The fewer models they serve, the less excess capacity they need to provision. Serving more models also adds a significant ops burden on the SREs and trust& safety teams.
- alex1138 15d agoIt's a shame Google crams it ham-fistedly into search results and that Google has some of the reputation it has because I actually really enjoy Gemini and I don't even use it for the reason people often list which is that you can cross-reference it to stuff in your Google account
- aff-vasileva 15d agoThe model seems fast enough to solve your problem before Google finishes explaining which of its three products you need to open to access it.
- gere 15d agoI have mixed feelings about Gemini 3.7 Flash. I used it for a personal project in Java and it was ok: it was crazy fast and it reached the correct result, but the code quality was barely passable. I also used it for a an app for my Garmin watch, and it wasn't good. The code was compiling, but functionality was totally broken and even with a lot of steering it wasn't able to make it work. GLM 5.3-flash instead was up for it and the code wasn't bad at all. I am curious to see if 3.8 is an improvement in this use case.
- koalaman 15d agoI use Gemini to make sense of things Claude says to me.
- EFLKumo 15d agoSomething maybe unfamiliar with you: not about coding but writing. I've asked it to write an argumentative essay, which is a part of "gaokao" (China's university entrance exam), and its work is *extremely* impressive. speaks and writes like a real senior high school student, and the opinions unfold progressively with deep hierarchy. I don't know how the Gemini team reaches this because this kind of Chinese capability literally outperforms at least 2/3 Chinese students, no to mention those who speak Chinese. After all, the model speaks like a real humankind if you prompt it well. That's AGI guys
- cubefox 14d agoNitpick, but in my opinion an LLM is an "it", not a "her" or "he". Using male or female pronouns risks anthropomorphizing them which can lead to unhealthy outcomes.
- EFLKumo 14d agoSorry! I was just a bit excited writing the comment and ignored that :(
- cubefox 14d agoNo worries.
- adleyjulian 14d agoFYI in Chinese he/she/it all use the same pronoun "ta" when spoken.
- SchemaLoad 14d agoThey do have a separate character for animals and objects 它 vs 他/她, though I can imagine learning english and just learning ta = he.
- livinglist 14d ago
- therealmarv 14d agoOn my short tests: This model is amazing and the speed makes it feel like another sort of AI. But it's bad at code reviews (maybe it's the harness agy cli?). Could not get it to same quality level on reviews like Opus, GPT 5.6, Grok. Even tried special code review skills but no luck.
- sreekanth850 14d agoDear Google, Kindly make you chat window on the right side of vscode in antigravity extension, There is a reason others kept it like that. I can see the code and inspect the files changed while Agents keep working. its critical for me personally.
- kelvinjps10 14d agoI think about Google is the value you get of their plans, for 5$ a month you get their ai plus model combined with 400gb you can share this with your family. The other ai companies don't provide family plans
- Sir_Twist 14d agoAnd the free year-long trial for college students they recently offered, which includes 5 tb of Google Drive storage.
- kelvinjps10 14d agoI got 6months for free when I bought my s25.
- brap 14d agoPeople have been sleeping on Gemini lately but these last few Flash releases (which were very rapid) are damn good. These sort of fast and cheap models are great for tasks that are verifiable and can be retried infinitely (like coding), you can basically get frontier results with a good harness (at a fraction of the time and money).
- mvdtnz 14d agoAs someone who has stubbornly stuck with Claude Code, what's a good harness for Gemini models?
- brap 14d agoAntigravity has been also rapidly improving lately, and your can also use any of the open coding harnesses. But I mostly meant “harness” as in your workflow/loop setup.
- _aavaa_ 14d agoDo they officially support you using your subscription in other harnesses?
- watusername 14d agoNo, and Google actively bans people for using their subscription from other harnesses via various proxies/gateways. To preempt certain replies, yes, I know you can pay API prices and use whatever harness you want.
- nharada 14d agoMeanwhile I pay for Pro and still don't have access to 3.7?
- almog 14d agoSame for me (at least through the Gemini app).
- im_soul 14d agodisclaimer : Introductory price expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply.
- mrbonner 14d agoI’m interested in a general knowledge model (closed or open weight) and not coding specific. I want to plan for travel and trip. Do you have one of your favorite HN crowd?
- abixb 14d agoI like Google's strategy here. These new Flash models of late (Flash 3.6, 3.7 and now 3.8) have obviously been distilled from a much larger unreleased model (Gemini 3.5 Pro, iirc from the rumors). One aspect of model releases that don't get discussed as much are the cache invalidation (changes in underlying architecture, weights, or tokenizers); I assess Google seems to be squeezing the maximum out of the last 'Pro' version they released with 3.1 back in February. Small models cataching up with their bigger siblings are fantastic news.
- alephnerd 14d agoA couple larger GCP customers requested this for sometime, especially on the cybersecurity side. A SOC/IR or AppSec team doesn't need a generalized model that knows when Chaucer lived but it absolutely needs a model that can efficiently, quickly, and accurately prioritize vulnerability severity or validate patches.
- 2001zhaozhao 14d agoHow generous is the Google subscription quotas compared to Anthropic and OpenAI? This sounds like a really good potential model for high volume due to its speed and cost effectiveness. (By high volume I mean things like "main app just updated with XYZ commits, please scan XYZ plugins and surface any compatibility issues")
- 654wak654 14d agoI'm on the Ultra plan and use it for chat, antigravity, and some other work automations (similar to your example). The only time I've ever hit my limit is when I use Deep Think (which usually eats up 4-5% of the 6-hour usage limit per response).
- thereitgoes456 14d agoReally generous. I'm on the Pro plan and I just use Antigravity for vibe coding w/o automation. It's actually difficult to hit my weekly limit now, it takes about ~30-35 hours of continuous agent work, which virtually only happens when building a new app from scratch.
- levelZero 14d agoGemini 3.8 flash thinks Entoloma sinuatum is good to eat... Otherwise feels great
- DrewKeller 14d ago[dead]
- yipinwong 14d agoAs a big proponent of GPT-5.6-Luna for the combination of speed/perf/(especially)price, Flash 3.8 seems like where I can specify Flash3.8 as the coding model as part of agent workflow. The video recognition is especially impressive as they got all of Youtube to train from. - Def people who has to queue video recognition jobs to use the model.
- lysecret 14d agoAlso just want to let my appreciation here for 3.7 it’s cheap super fast super reliable incredible at information parsing eu host able (important for us) and perfectly integrated into gcp. Great job google!
- MrBuddyCasino 14d agoI hope they bring a lite version, its good enough for information parsing and very cheap.
- weird-eye-issue 14d agoUse Luna for that
- npn 14d agoStill refuse to search internet for stuff it thinks does not exist lol. And even when searching for internet, it still cannot suggest a up-to-date approach to the problem. For example I'm using crystal, it recently revamped the concurrency/parallel model. Even using web search, gemini still does not aware of the new feature and still give the outdated code. I'm sure my crystal usage is not the unique case here.
- luciana1u 14d ago[flagged]
- HardCodedBias 14d agoI have to say: The Google brand remains powerful on HN! I’m shocked.
- henry-xli 14d agoI can’t wait until waiting hours and spending a big chunk of your usage per task seems antiquated, and real-time iteration on massive code changes is the norm. This might just be the year of efficiency, that truly allows AI to be used to the heart’s content.
- mark_l_watson 14d agoWhen deepseek-v4-flash-0731 was released I used it constantly for everything, and I loved the very low cost and speed. I think Google has the same game plan with their flash models.
- ddp26 14d agoThere must be a deeper read on why Google can rapidly ship better small models while being delayed months on the bigger model. What's the simplest explanation?
- cogman10 14d agoPerhaps post training? I believe I read that Qwen 3.8 is just post trained Qwen 3.6, which is why it was able to be released so quick. It may be that these flash models are simply post trained larger older models.
- firemelt 14d agoI wish google to thrive
- Helldez 14d ago[dead]
- jetter 14d agoCAD for 3D printing is finally becoming feasible with Flash 3.7 and 3.8. Exciting times. https://github.com/ModelRift/openscad-skill/ https://github.com/ModelRift/openscad-skill/
- ldm0 14d agoIt’s strange that its score on Terminal‑Bench 4.0 is so low. They aren’t fast enough to benchmaxx that section.
- pampas 14d agoGemini 3.8 Flash is top of the Redactle LLM benchmark but so was Gemini 3.7 Flash. Both one shot all puzzles in the evals though 3.8 is just a bit faster. It also does the evals cheaper and faster than almost all the other models I've tried. https://redactle.net/llm-leaderboard https://redactle.net/llm-leaderboard
- centaurz 14d agoA company with 400+B revenue from software cannot build a usable command line cli for its vital AI model?
- japgolly 14d agoThey do have a cli: https://github.com/google-antigravity/antigravity-cli https://github.com/google-antigravity/antigravity-cli
- jpau 14d agoThe iteration cycle is becoming very quick. Gemini 3.8 Flash arrived just 20 days after 3.7 Flash. Similarly Qwen3.8-Max was updated in just 30 days (to the 0902 release) and Muse Spark in just 28 days (to the 1.3 release). A year ago iterative releases were every 3-6 months. At what point will they reach nightly candidates?
- mohamedkoubaa 14d agoThe race to the bottom continues
- Galorious 14d agoIs anyone here using using these models via google subscription (not api). I tried to in the past using gemini cli and then agy - headless invoked by codex and claude code, but they were so incredibly buggy that it stalled 1/2 times and I cancelled. Interested to know if that has changed!
- zuzululu 14d agonot really getting the excitement over this, its at opus 5 medium level, and opus 5 is not really the go to model , claude purists hate it so its fast sure and decent at non coding usage but for developers nothing can really top sol or fable. even grok 4.6 is so so and i would not choose 3.8 flash over it.
- Razengan 14d agoWhat is with Google's dumb ass STILL refusing to respect the OS dark mode setting in fucking 2027??
- schmorptron 14d agoA reminder that google is the only major lab without a meaningful opt-out of training on your data. The only way to opt out is to disable message history entirely, which seems like a darkest of dark patterns to get users to leave "opt in" to training on, because next to nobody wants to use it without message history.
- newppc 14d agoIf Google has the juice and wants to win, they need to start releasing world models.
- drivebyhooting 14d agoI’ve used the Gemini flash, but then when I have soul ultra check its work, it found a bunch of cut corners and improper design. As much as I like the speed and interactivity, I really don’t trust it
- johnnyApplePRNG 14d agoWhy is it still such a bad coding agent? Does anybody have any insight? I am continually impressed with Gemini's chat responses, which encourages me to test their agentic capabilities and... no... no... and no... every single time. It's terrifying watching it, really.
- robertwt7 14d agothis is cool for all other non coding task. however I am still stuck on 3.6 flash on my gemini web as a plus user, can anyone else even access 3.7 flash in AU?
- alvah 14d agoAU Pro user here. 3.8 Flash available (default) in the web app for me.
- Alpha3031 14d agoAU user also, just checked AI studio since that seemed like the best bet and both 3.8 and 3.7 show up (and can be used for chat in playground, though IDK what the limits for that are). Chat in gemini.google.com is also 3.6 for me but I'm on free tier lol so I don't exactly expect it to show up any time soon. I think there's also another free API beyond the AI studio one (which is 20 RPD free according to docs so not really useful) but I forgot where it was (Google cloud maybe?) and what the limits for that were.
- akurilin 14d agoCurious which model this can supplant as a clear winner on almost every metric. Sol? Looks like it's not quite there on a couple of benches, but I'm not clear how much they matter in practice.
- throw10920 14d agoWe've gotten an unusually fast speed of Gemini Flash releases over the past few months. Is this Recursive Self Improvement, or Google just trying to distract from the fact that it's been a while since the last Gemini Pro release?
- tkgally 14d agoThe blog post says it is RSI: “both of today's releases are … accelerated by long-running agentic loops designed to recursively evaluate and refine the underlying models.”
- throw10920 14d agoYeah, but Google is incentivized to claim that regardless of truth value. Critical analysis is necessary.
- tagalog 14d agoGemini flash seems to have been a bit of a sleeper. Somehow it's ended up as the most used LLM for my client document extraction work these past few months. I have an eval harness that runs every Thursday to determine which models are the current best for a few different client workflows. And since May(?) flash has slowly been taking over more and more stuff to the point it is now 100% on 8 out of 11 document extraction flows with the other 3 being a Flash / Opus 4.8 mix for high value stuff where cost is less of a factor.
- thrdbndndn 14d agoWhen can we use it in Gemini (web)? It still uses 3.6 Flash for example.
- xyzkoi 14d ago[dead]
- maxnevermind 14d agoJust tried Gemini 3.8 Flash on these 2 consecutive prompts at gemini.google.com: 1 what is tesla cybercab plan to address legal implications of accident that will happen? who is going to be responsible for them when they happen? are they covered by tesla insurance or some other insurance? are there any official plan/statements around that? 2 what was the name of the experiment they started in san antonio tx when some cars didn't have a driver? what was the results of it? did they expand the operations? it was much smaller than waymo, is it growing? how it is related to robotaxi? It is not able to connect the dots that I keep asking about Tesla in 2nd prompt and spit out some unrelated stuff. Really? How it can be that bad? Gemini 3.1 Pro model works fine in this case btw. I thought maybe it is about knowledge cut over date and it doesn't know about those events from 2025 but it seems it has the knowledge up to March 2025. Top 10 in Intelligence on artificialanalysis ladies and gentlemen.
- casey2 14d agoMeh, not any noticeable improvement and unlike 3.7 high it eats all your credits, perhaps medium would be better
- asdaqopqkq 14d agoGemini models look so good on paper by IRL dev and daily life usage totally make it seem like it's way behind Codex and Claude.
- yoga666 14d ago[flagged]
- 1saadcodes 14d agoThe recent Sonnet models have been disappointing for me personally which is why I'm going look into using Opus/Fable as the planner and Flash as the executor. Let the expensive model handle the hard thinking and use Flash for implementation and tests so that I can stretch the Opus/Fable usage further
- virajk_31 14d agoGemini 3.7 benchmarks against GDPVal-AA-V2 were 1525 in Aug blog post. However same model against same benchmark is 1482 in today's blog post of Gemini 3.8 release.. Do they make it intentionally to look previous model less superior than current models? or these are the real numbers when re-ran the benchmark..?
- PaulStatezny 14d agoIt struck me today using Google Antigravity (Claude Code alternative) just how direct and usably terse Gemini is in prose. I've complained plenty on here about Claude verbosity and TED-talk phrasing, and it seems by contrast Gemini has already arrived at the dream end-state of Claude from a prose standpoint. Sometimes I ask for feedback, and I get back a list of multiple-choice options as if it's already ready to go. If I indicate I'm thinking about doing something, sometimes it'll just...do it. (Not in an annoying way.) It seems very geared toward action in a way that's completely refreshing coming from months steeped in Claude essays.
- nxdmum 14d agoso far using Gemini from 2.5 Pro to date (3.7 flash) - the way google trains the model it seems - is to identify top 3 to 4 things to fix first. as a result gemini is not that good in being thorough - but it's a needle mover . Opus5 Opus4.8 and Fable always point out things that Gemini missed. but Gemini was a needle mover - identifying the most important things to fix. I always enjoy interacting with Gemini . When i ask it questions about designing a new model etc - it's always the most helpful and encouraging . I really want to thank the Google team for this and their happy positive models they generate. I use gemini flash after a round of deliberation between Sol and Kimi these days on the main plan . Kimi 2.7 paired with Gemini 3.5/4.6/3.7 flash has been my implementer - Kimi k3 and Sol 5.6 have been my planners and code reviewers. I would have ideally like Anthropic and was on their USD 200 plan - but after they didnt sign the letter for Open source models - i dropped my subscription . Wont make a difference to their lives. But Sol is great . Combined with Kimi K3 for adversarial plan reviews - you get robust plans . And Sol as a reviewer for Gemini/kimi 2.7 - you get great edge case handling and robust code. I also integrated muse 1.2 - on the contributor tier - and I Just saw facebook release 1.3 muse . this is great news. The training im permitting is my thanks to FB for releasing the open source models of the past ! Thank you !
- tabs_or_spaces 14d agoIt's really disappointing to see social media dismissing gemini so easily. I think the worst thing we can do is have loyalty towards models. I used to be loyal towards Claude, and my viewpoint changed dramatically when I used codex. I highly recommend that if you are someone who only used one model so far, that you really give another model a shot and see how it goes. It's very eye opening and gives you a more holistic perspective. Vendor locking is a big problem when it comes to models, and I hope the software world doesn't do this blindly.
- szundi 14d ago[dead]
- anilakar 14d agoWill Gemini refuse to work like Claude when it hits modulo 11 calculation on a number string or finds a variable named CVV?
- dongking 14d agoThe speed, combined with the fact that this thing is really good at JavaScript, is pretty exciting. I’ve added another AI programming assistant to my toolkit; hopefully AI will continue to get stronger.
- or1gaminal 14d agoI'm running lexical analysis on gemini 3.8 flash and the latency progress is incredible. For my tasks the latency is reduced by ~40% w/ quality on par.
- greenjudge 14d ago[flagged]
- k9294 14d agoIs there any comparison of usage limits for Antigravity plans vs. Codex? I just ran two light tasks on my codebase and got 100% of the weekly limits of a Pro plan blown away. Is Ultra plan any different? Because on Codex it wouldn't affect my Max plan at all, I think it would have been below 1% othese usage.
- mark_l_watson 14d agoI am not surprised. Google is a real company that wants to make money on selling services. Anthropic and OpenAI are in a different game of spending investor money to buy market share.
- XCSme 14d agoIn my tests 3.8 Flash is considerably more expensive[0]/less token efficient than 3.7 or 3.6, and not necessarily much smarter. I assume it is faster in tps, but hard to tell because ot also outputs more tokens, so response time is slower oferall. [0]: https://aibenchy.com/compare/google-gemini-3-6-flash-high/google-gemini-3-7-flash-high/google-gemini-3-8-flash-high/ https://aibenchy.com/compare/google-gemini-3-6-flash-high/go...
- theplumber 14d agoAt this point I wonder why Google still releases proprietary models. They are behind the open models.
- Unified-Mentor 14d ago[dead]
- hn4e309qn4 14d ago[dead]
- Surac 14d agoI think hackernews should not suport closed moddels
- albrewer 14d agoIf I get a better Gemma 26b-a4b MoE model out of it I'm all for it. Really wishing Qwen 3.8 would release a MoE variant of their 27b model size, but I wont complain if google beats them to it.
- Foobar8568 14d agoWhere is my Gemma 5?
- spwa4 14d agoAlexandr Wang (Meta's new AI chief) makes a good summary of this model: https://x.com/alexandr_wang/status/2079707749412483104?lang=en https://x.com/alexandr_wang/status/2079707749412483104?lang=...
- xpuente 14d agoHe inherited his predecessor’s bad manners and arrogance.
- sriniwasx 14d ago[dead]
- keytalker 14d agocongrats to the gemini team, the work is super impressive! I switched my subagent-swarm skill(https://github.com/bazelment/yoloswe/blob/main/.claude/skills/subagent-swarm/SKILL.md https://github.com/bazelment/yoloswe/blob/main/.claude/skill...) to the 3.8 model and has been very happy so far. It has been driving the swarm to produce steady outcome, and more importantly, the exec communication is also crystal clear, instead of filling with jargons and long sentences.
- _s_a_m_ 13d agothe names are getting so ridiculous and childish, are we here in ideocracy?