32 ms·
Gemini 3
https://blog.google/technology/developers/gemini-3-developers/ https://blog.google/technology/developers/gemini-3-developer...
https://aistudio.google.com/prompts/new_chat?model=gemini-3-pro-preview https://aistudio.google.com/prompts/new_chat?model=gemini-3-...
- nilsingwersen 10mo agoFeeling great to see something confidential
- RobinL 10mo ago- Anyone have any idea why it says 'confidential'? - Anyone actually able to use it? I get 'You've reached your rate limit. Please try again later'. (That said, I don't have a paid plan, but I've always had pretty much unlimited access to 2.5 pro) [Edit: working for me now in ai studio]
- sd9 10mo agoHow long does it typically take after this to become available on https://gemini.google.com/app https://gemini.google.com/app ? I would like to try the model, wondering if it's worth setting up billing or waiting. At the moment trying to use it in AI Studio (on the Free tier) just gives me "Failed to generate content, quota exceeded: you have reached the limit of requests today for this model. Please try again tomorrow."
- Squarex 10mo agoToday I guess. They were not releasing the preview models this time and it seems the want to synchronize the release.
- mpeg 10mo agoAllegedly it's already available in stealth mode if you choose the "canvas" tool and 2.5. I don't know how true that is, but it is indeed pumping out some really impressive one shot code Edit: Now that I have access to Gemini 3 preview, I've compared the results of the same one shot prompts on the gemini app's 2.5 canvas vs 3 AI studio and they're very similar. I think the rumor of a stealth launch might be true.
- sd9 10mo agoThanks for the hint about Canvas/2.5. I have access to 3.0 in AI Studio now, and I agree the results are very similar.
- csomar 10mo agoIt's already available. I asked it "how smart are you really?" and it gave me the same ai garbage template that's now very common on blog posts: https://gist.githubusercontent.com/omarabid/a7e564f09401a64e7daf7f5f84351e46/raw/9024dd6ca946ac442d0a5fed81da2d56fd72c5db/gistfile1.txt https://gist.githubusercontent.com/omarabid/a7e564f09401a64e...
- magicalhippo 10mo ago> https://gemini.google.com/app https://gemini.google.com/app How come I can't even see prices without logging in... they doing regional pricing?
- Romario77 10mo agoIt's available in cursor. Should be there pretty soon as well.
- ionwake 10mo agoare you sure its available in cursor? ( I get: We're having trouble connecting to the model provider. This might be temporary - please try again in a moment. )
- netdur 10mo agoOn gemini.google.com, I see options labeled 'Fast' and 'Thinking.' The 'Thinking' option uses Gemini 3 Pro
- mil22 10mo agoIt's available to be selected, but the quota does not seem to have been enabled just yet. "Failed to generate content, quota exceeded: you have reached the limit of requests today for this model. Please try again tomorrow." "You've reached your rate limit. Please try again later." Update: as of 3:33 PM UTC, Tuesday, November 18, 2025, it seems to be enabled.
- misiti3780 10mo agoseeing the same issue.
- sottol 10mo agoyou can bring your google api key to try it out, and google used to give $300 free when signing up for billing and creating a key. when i signed up for billing via cloud console and entered my credit card, i got $300 "free credits". i haven't thrown a difficult problem at gemini 3 pro it yet, but i'm sure i got to see it in some of the A/B tests in aistudio for a while. i could not tell which model was clearly better, one was always more succinct and i liked its "style" but they usually offered about the same solution.
- lousken 10mo agoI hope some users will switch from cerebras to free up those resources
- sarreph 10mo agoLooks to be available in Vertex. I reckon it's an API key thing... you can more explicitly select a "paid API key" in AI Studio now.
- CjHuber 10mo agoFor me it’s up and running. I was doing some work with AI Studio when it was released and reran a few prompts already. Interesting also that you can now set thinking level low or high. I hope it does something, in 2.5 increasing maximum thought tokens never made it think more
- 10mo ago
- guluarte 10mo agoit is live in the api > gemini-3-pro-preview-ais-applets > gemini-3-pro-preview
- spudlyo 10mo agoCan confirm. I was able to access it using GPTel in Emacs using 'gemini-3-pro-preview' as the model name.
- informal007 10mo agoIt seem that Google doesn't prepare well to release Gemini 3 but leak many contents, include the model card early today and gemini 3 on aistudio.google.com
- __jl__ 10mo agoAPI pricing is up to $2/M for input and $12/M for output For comparison: Gemini 2.5 Pro was $1.25/M for input and $10/M for output Gemini 1.5 Pro was $1.25/M for input and $5/M for output
- raincole 10mo agoStill cheaper than Sonnet 4.5: $3/M for input and $15/M for output.
- brianjking 10mo agoIt is so impressive that Anthropic has been able to maintain this pricing still.
- Aeolun 10mo agoBecause every time I try to move away I realize there’s nothing equivalent to move to.
- Alex-Programs 10mo agoPeople insist upon Codex, but it takes ages and has an absolutely hideous lack of taste.
- DeathArrow 10mo agoIt generated a quite cool pelican on a bike: https://imgur.com/a/yzXpEEh https://imgur.com/a/yzXpEEh
- rixed 10mo ago2025: solve the biking pelican problem 2026: cure cancer
- GodelNumbering 10mo agoAnd of course they hiked the API prices Standard Context(≤ 200K tokens) Input $2.00 vs $1.25 (Gemini 3 pro input is 60% more expensive vs 2.5) Output $12.00 vs $10.00 (Gemini 3 pro output is 20% more expensive vs 2.5) Long Context(> 200K tokens) Input $4.00 vs $2.50 (same +60%) Output $18.00 vs $15.00 (same +20%)
- CjHuber 10mo agoIs it the first time long context has separate pricing? I hadn’t encountered that yet
- Topfi 10mo agoGoogle has been doing that for a while.
- brianjking 10mo agoGoogle has always done this.
- CjHuber 10mo agoOk wow then I‘ve always overlooked that.
- 1ucky 10mo agoAnthropic is also doing this for long context >= 200k Tokens on Sonnet 4.5
- deleted 10mo ago[deleted]
- panarky 10mo agoClaude Opus is $15 input, $75 output.
- xnx 10mo agoIf the model solves your needs in fewer prompts, it costs less.
- aliljet 10mo agoWhen will this be available in the cli?
- _ryanjsalva 10mo agoGemini CLI team member here. We'll start rolling out today.
- Sammi 10mo agoI'm already seeing it in https://aistudio.google.com/ https://aistudio.google.com/
- deleted 10mo ago[deleted]
- skerit 10mo agoNot the preview crap again. Haven't they tested it enough? When will it be available in Gemini-CLI?
- CjHuber 10mo agoHonestly I liked 2.5 Pro preview much more than the final version
- prodigycorp 10mo agoI'm sure this is a very impressive model, but gemini-3-pro-preview is failing spectacularly at my fairly basic python benchmark. In fact, gemini-2.5-pro gets a lot closer (but is still wrong). For reference: gpt-5.1-thinking passes, gpt-5.1-instant fails, gpt-5-thinking fails, gpt-5-instant fails, sonnet-4.5 passes, opus-4.1 passes (lesser claude models fail). This is a reminder that benchmarks are meaningless – you should always curate your own out-of-sample benchmarks. A lot of people are going to say "wow, look how much they jumped in x, y, and z benchmark" and start to make some extrapolation about society, and what this means for others. Meanwhile.. I'm still wondering how they're still getting this problem wrong. edit: I've a lot of good feedback here. I think there are ways I can improve my benchmark.
- m00dy 10mo agothat's why everyone using AI for code should code in rust only.
- Filligree 10mo agoWhat's the benchmark?
- prodigycorp 10mo agonice try!
- ankit219 10mo agoyou already sent the prompt to gemini api - and they likely recorded it. So in a way they can access it anyway. Posting here or not would not matter in that aspect.
- petters 10mo agoGood personal benchmarks should be kept secret :)
- mlrtime 10mo ago
- nickandbro 10mo agoWhat we have all been waiting for: "Create me a SVG of a pelican riding on a bicycle" https://www.svgviewer.dev/s/FfhmhTK1 https://www.svgviewer.dev/s/FfhmhTK1
- Thev00d00 10mo agoThat is pretty impressive. So impressive it makes you wonder if someone has noticed it being used a benchmark prompt.
- burkaman 10mo agoSimon says if he gets a suspiciously good result he'll just try a bunch of other absurd animal/vehicle combinations to see if they trained a special case: https://simonwillison.net/2025/Nov/13/training-for-pelicans-riding-bicycles/ https://simonwillison.net/2025/Nov/13/training-for-pelicans-...
- jmmcd 10mo ago"Pelican on bicycle" is one special case, but the problem (and the interesting point) is that with LLMs, they are always generalising. If a lab focussed specially on pelicans on bicycles, they would as a by-product improve performance on, say, tigers on rollercoasters. This is new and counter-intuitive to most ML/AI people.
- BoorishBears 10mo agoThe gold standard for cheating on a benchmark is SFT and ignoring memorization. That's why the standard for quickly testing for benchmark contamination has always been to switch out specifics of the task. Like replacing named concepts with nonsense words in reasoning benchmarks.
- jmmcd 10mo agoYes. But "the gold standard" just means "the most natural, easy and dumb way".
- CjHuber 10mo agoInteresting that they added an option to select your own API key right in AI studio‘s input field. I sincerely hope the times of generous free AIstudio usage are not over
- golfer 10mo agoSupposedly this is the model card. Very impressive results. https://pbs.twimg.com/media/G6CFG6jXAAA1p0I?format=jpg&name=medium https://pbs.twimg.com/media/G6CFG6jXAAA1p0I?format=jpg&name=... Also, the full document: https://archive.org/details/gemini-3-pro-model-card/page/n3/mode/2up https://archive.org/details/gemini-3-pro-model-card/page/n3/...
- tweakimp 10mo agoEvery time I see a table like this numbers go up. Can someone explain what this actually means? Is there just an improvement that some tests are solved in a better way or is this a breakthrough and this model can do something that all others can not?
- rvnx 10mo agoThis is a list of questions and answers that was created by different people. The questions AND the answers are public. If the LLM manages through reasoning OR memory to repeat back the answer then they win. The scores represent the % of correct answers they recalled.
- tylervigen 10mo agoThat is not entirely true. At least some of these tests (like HLE and ARC) take steps to keep the evaluation set private so that LLMs can’t just memorize the answers. You could question how well this works, but it’s not like the answers are just hanging out on the public internet.
- samuelknight 10mo ago"Gemini 3 Pro Preview" is in Vertex
- ponyous 10mo agoCan’t wait to test it out. Been running a tons of benchmarks (1000+ generations) for my AI to CAD model project and noticed: - GPT-5 medium is the best - GPT-5.1 falls right between Gemini 2.5 Pro and GPT-5 but it’s quite a bit faster Really wonder how well Gemini 3 will perform
- santhoshr 10mo agoPelican riding a bicycle: https://pasteboard.co/CjJ7Xxftljzp.png https://pasteboard.co/CjJ7Xxftljzp.png
- mohsen1 10mo agoSome time I think I should spend $50 on Upwork to get a real human artist to do it first to know what is that we're going for. What a good pelican riding a bicycle SVG is actually looking like?
- AstroBen 10mo agoIMO it's not about art, but a completely different path than all these images are going down. The pelican needs tools to ride the bike, or a modified bike. Maybe a recumbent?
- robterrell 10mo agoAt this point I'm surprised they haven't been training on thousands of professionally-created SVGs of pelicans on bicycles.
- notatoad 10mo agoi think anything that makes it clear they've done that would be a lot worse PR than failing the pelican test would ever be.
- imiric 10mo agoIt would be next to impossible for anyone without insider knowledge to prove that to be the case. Secondly, benchmarks are public data, and these models are trained on such large amounts of it that it would be impractical to ensure that some benchmark data is not part of the training set. And even if it's not, it would be safe to assume that engineers building these models would test their performance on all kinds of benchmarks, and tweak them accordingly. This happens all the time in other industries as well. So the pelican riding a bicycle test is interesting, but it's not a performance indicator at this point.
- Der_Einzige 10mo agoWhen will they allow us to use modern LLM samplers like min_p, or even better samplers like top N sigma, or P-less decoding? They are provably SOTA and in some cases enable infinite temperature. Temperature continues to be gated to maximum of 0.2, and there's still the hidden top_k of 64 that you can't turn off. I love the google AI studio, but I hate it too for not enabling a whole host of advanced features. So many mixed feelings, so many unanswered questions, so many frustrating UI decisions on a tool that is ostensibly aimed at prosumers...
- ttul 10mo agoMy favorite benchmark is to analyze a very long audio file recording of a management meeting and produce very good notes along with a transcript labeling all the speakers. 2.5 was decently good at generating the summary, but it was terrible at labeling speakers. 3.0 has so far absolutely nailed speaker labeling.
- iagooar 10mo agoWhat prompt do you use for that?
- gregsadetsky 10mo agoI just tried "analyze this audio file recording of a meeting and notes along with a transcript labeling all the speakers" (using the language from the parent's comment) and indeed Gemini 3 was significantly better than 2.5 Pro. 3 created a great "Executive Summary", identified the speakers' names, and then gave me a second by second transcript: [00:00] Greg: Hello. [00:01] X: You great? [00:02] Greg: Hi. [00:03] X: I'm X. [00:04] Y: I'm Y. ... Super impressive!
- HPsquared 10mo agoDoes it deduce everyone's name?
- gregsadetsky 10mo agoIt does! I redacted them, but yes. This was a 3-person call.
- punnerud 10mo agoI made a simple webpage to grab text from YouTube videos: https://summynews.com https://summynews.com Great for this kind of testing? (want to expand to other sources in the long run)
- valtism 10mo ago
- denysvitali 10mo agoFinally!
- thedelanyo 10mo agoReading the introductory passage - all I can say now is, Ai is here to stay.
- meetpateltech 10mo agoDeepMind page: https://deepmind.google/models/gemini/ https://deepmind.google/models/gemini/ Gemini 3 Pro DeepMind Page: https://deepmind.google/models/gemini/pro/ https://deepmind.google/models/gemini/pro/ Developer blog: https://blog.google/technology/developers/gemini-3-developers/ https://blog.google/technology/developers/gemini-3-developer... Gemini 3 Docs: https://ai.google.dev/gemini-api/docs/gemini-3 https://ai.google.dev/gemini-api/docs/gemini-3 Google Antigravity: https://antigravity.google/ https://antigravity.google/
- fuzzythinker 10mo agoAlso recently: Code Wiki: https://codewiki.google/ https://codewiki.google/
- wohoef 10mo agoCurious to see it in action. Gemini 2.5 has already been very impressive as a study buddy for courses like set theory, information theory, and automata. Although I’m always a bit skeptical of these benchmarks. Seems quite unlikely that all of the questions remain out of their training data.
- fosterfriends 10mo agoGemini 3 and 3 pro are good bit cheaper than Sonnet 4.5 as well. Big fan
- aliljet 10mo agoUnderstanding precisely why Gemini 3 isn't front of the pack on SWE Bench is really what I was hoping to understand here. Especially for a blog post targeted at software developers...
- cube2222 10mo agoYeah, they mention a benchmark I'm seeing the first time (Terminal-Bench 2.0) and are supposedly leading in, while for some reason SWE Bench is down from Sonnet 4.5. Curious to see some third-party testing of this model. Currently it seems to primarily improve of "general non-coding and visual reasoning" primarily, based on the benchmarks.
- nico1207 10mo agoThey are not even leading in Terminal-Bench... GPT 5.1-codex is better than Gemini 3 Pro
- pawelduda 10mo agoWhy is this particular benchmark important?
- aliljet 10mo agoThus far, this is one of the best objective evaluations of real world software engineering...
- adastra22 10mo agoIdk, Sonnet 4.5 score better than Sonnet 4.0 on that benchmark, but is markedly worse in my usage. The utility of the benchmark is fading as it is gamed.
- pertymcpert 10mo agoI find 4.5 a much better model FWIW.
- svantana 10mo agoGrok got to hold the top spot of LMArena-text for all of ~24 hours, good for them [1]. With stylecontrol enabled, that is. Without stylecontrol, gemini held the fort. [1] https://lmarena.ai/leaderboard/text https://lmarena.ai/leaderboard/text
- inkysigma 10mo agoIs it just me or is that link broken because of the cloudflare outage? Edit: nvm it looks to be up for me again
- dyauspitr 10mo agoGrok is heavily censored though
- KingMob 10mo agoIs it censored... or just biased towards edge-lord MechaHitler nonsense whenever Musk feels like tinkering with the system prompt?
- tempaccountabcd 10mo ago[dead]
- bnchrch 10mo agoI've been so happy to see Google wake up. Many can point to a long history of killed products and soured opinions but you can't deny theyve been the great balancing force (often for good) in the industry. - Gmail vs Outlook - Drive vs Word - Android vs iOS - Worklife balance and high pay vs the low salary grind of before. Theyve done heaps for the industry. Im glad to see signs of life. Particularly in their P/E which was unjustly low for awhile.
- digbybk 10mo agoIronically, OpenAI was conceived as a way to balance Google's dominance in AI.
- dragonwriter 10mo agoI thought it was a workaround to Google's complete disinterest in productizing the AI research it was doing and publishing, rather than a way to balance their dominance in a market which didn't meaningfully exist.
- mattnewton 10mo agoThat’s how it turned out, but IIRC at the time of OpenAI’s founding, “AI” was search and RL which Google and deep mind were dominating, and self driving, which Waymo was leading. And OpenAI was conceptualized as a research org to compete. A lot has changed and OpenAI has been good at seeing around those corners.
- jpadkins 10mo agoElon Musk specifically gave OAI $150M early on because of the risk of Google being the only Corp that has AGI or super-intelligence. These emails were part of the record in the lawsuit.
- jonny_eh 10mo agoThat was actually Character.ai's founding story. Two researchers at Google that were frustrated by a lack of resources and the inability to launch an LLM based chatbot. The founders are now back at Google. OpenAI was founded based on fears that Google would completely own AI in the future.
- recitedropper 10mo ago[flagged]
- 63stack 10mo agoI noticed this as well, you are already downvoted into gray
- recitedropper 10mo agoInevitable... certainly more so than AGI :)
- wasabi991011 10mo agoThey're downboted into grey because it's complaining about the future of this thread before it has even happened. Also it's conspiratorial, without much evidence.
- recitedropper 10mo agoPeek the other threads.
- hobofan 10mo agoAnd now it's flagged. I think this is one of HNs biggest weaknesses. If you are a sufficiently large engineering organization with enough employees that pass the self-moderation karma thresholds, you can essentially strike down any significantly critical discussion.
- zamadatix 10mo agoWithout a public moderation log (i.e. even user flags being part of the log) claims like this will always come up but to me it always seems more likely just the early commenting users tired of being told they are part of some astroturf campaign and if they don't flock to agree with the OPs views it must just be more proof. I'm sure both reasons happen to some degree, just as a matter of how often is actual astroturfing vs "a small percentage of active people can't possibly just have different thoughts than me".
- zurfer 10mo agoIt also tops LMSYS leaderboard across all categories. However knowledge cutoff is Jan 2025. I do wonder how long they have been pre-training this thing :D.
- mudkipdev 10mo agoIsn't it the same cutoff as 2.5?
- mihau 10mo ago@simonw wen pelican
- Workaccount2 10mo agoIt still failed my image identification test ([a photoshopped picture of a dog with 5 legs]...please count the legs) that so far every other model has failed agonizingly, even failing when I tell them they are failing, and they tend to fight back at me. Gemini 3 however, while still failing, at least recognized the 5th leg, but thought the dog was...well endowed. The 5th leg however is clearly a leg, despite being where you would expect the dogs member to be. I'll give it half credit for at least recognizing that there was something there. Still though, there is a lot of work that needs to be done on getting these models to properly "see" images.
- lukebechtel 10mo agoah interesting. I wonder if this is a "safety guardrails blindspot" due to the placement.
- achow 10mo agoOp is right. https://imgcdn.stablediffusionweb.com/2024/4/19/8e54925a-0048-4f5a-88d4-aef3eedfd93e.jpg https://imgcdn.stablediffusionweb.com/2024/4/19/8e54925a-004... For the above pic I asked "What is wrong with the image?" Result: - It totally missed the most obvious one - six fingers. Instead it said this: Based on a close inspection of the image, here are the things that are "wrong" or anomalous, indicating that it is likely an AI-generated image: - The Thumb Anatomy: This is the most obvious error. The digit in the thumb's position (far left) looks exactly like a long index finger. A natural human thumb has two segments (phalanges) and is shorter and thicker, opposing the other fingers. This digit appears to have three segments and is roughly the same length and shape as the other fingers, giving the hand the appearance of having five fingers and no thumb. - Unnatural Skin Texture: The skin is unnaturally smooth, waxy, and free of normal human imperfections like distinct pores, fine wrinkles, or hair. It has a "plastic" or "airbrushed" look typical of generative AI. - Palm Lines: The creases on the palm are vague and generic. They don't follow the distinct pattern of major flexion creases (like the heart, head, and life lines) found on a real human hand. - Wrist Connection: The connection between the hand and the wrist looks slightly distorted, with tendons that appear too prominent and smooth for the pose. [Edit: 3.0 is same as 2.5 - both answered almost identically]
- icyfox 10mo agoPretty happy the under 200k token pricing is staying in the same ballpark as Gemini 2.5 Pro: Input: $1.25 -> $2.00 (1M tokens) Output: $10.00 -> $12.00 Squeezes a bit more margin out of app layer companies, certainly, but there's a good chance that for tasks that really require a sota model it can be more than justified.
- rudedogg 10mo agoEvery recent release has bumped the pricing significantly. If I was building a product and my margins weren’t incredible I’d be concerned. The input price almost doubled with this one.
- icyfox 10mo agoI'm not sure how concerned people should be at the trend lines. If you're building a product that already works well, you shouldn't feel the need to upgrade to a larger parameter model. If your product doesn't work and the new architectures unlock performance that would let you have a feasible business, even a 2x on input tokens shouldn't be the dealbreaker. If we're paying more for a more petaflop heavy model, it makes sense that costs would go up. What really would concern me is if companies start ratcheting prices up for models with the same level of performance. My hope is raw hardware costs and OSS releases keep a lid on the margin pressure.
- deleted 10mo ago[deleted]
- gertrunde 10mo ago"AI Overviews now have 2 billion users every month." "Users"? Or people that get presented with it and ignore it?
- singhrac 10mo agoThey're a bit less bad than they used to be. I'm not exactly happy about what this means to incentives (and rewards) for doing research and writing good content, but sometimes I ask a dumb question out of curiosity and Google overview will give it to me (e.g. "what's in flower food?"). I don't need GPT 5.1 Thinking for that.
- recitedropper 10mo ago"Since then, it’s been incredible to see how much people love it. AI Overviews now have 2 billion users every month." Cringe. To get to 2 billion a month they must be counting anyone who sees an AI overview as a user. They should just go ahead and claim the "most quickly adopted product in history" as well.
- mNovak 10mo agoMaybe you ignore it, but Google has stated in the past that click-through rates with AI overviews are way down. To me, that implies the 'user' read the summary and got what they needed, such that they didn't feel the need to dig into a further site (ignoring whether that's a good thing or not). I'd be comfortable calling a 'user' anyone who clicked to expand the little summary. Not sure what else you'd call them.
- gertrunde 10mo agoYou're right, I'm probably being a little uncharitable! Normal users (i.e. not grumpy techies ;) ) probably just go with the flow rather than finding it irritating.
- rvz 10mo agoI expect almost no-one to read the Gemini 3 model card. But here is a damning excerpt from the early leaked model card from [0]: > The training dataset also includes: publicly available datasets that are readily downloadable; data obtained by crawlers; licensed data obtained via commercial licensing agreements; user data (i.e., data collected from users of Google products and services to train AI models, along with user interactions with the model) in accordance with Google’s relevant terms of service, privacy policy, service-specific policies, and pursuant to user controls, where appropriate; other datasets that Google acquires or generates in the course of its business operations, or directly from its workforce; and AI-generated synthetic data. So your Gmails are being read by Gemini and is being put on the training set for future models. Oh dear and Google is being sued over using Gemini for analyzing user's data which potentially includes Gmails by default. Where is the outrage? [0] https://web.archive.org/web/20251118111103/https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf https://web.archive.org/web/20251118111103/https://storage.g... [1] https://www.yahoo.com/news/articles/google-sued-over-gemini-ai-151508072.html https://www.yahoo.com/news/articles/google-sued-over-gemini-...
- aoeusnth1 10mo agoThis seems like a dubious conclusion. I think you missed this part: > in accordance with Google’s relevant terms of service, privacy policy
- stefs 10mo agoi'm very doubtful gmail mails are used to train the model by default, because emails contain private data and as soon as this private data shows up in the model output, gmail is done. "gmail being read by gemini" does NOT mean "gemini is trained on your private gmail correspondence". it can mean gemini loads your emails into a session context so it can answer questions about your mail, which is quite different.
- inkysigma 10mo agoIsn't Gmail covered under the Workspace privacy policy which forbids using that for training data. So I'm guessing that's excluded by the "in accordance" clause.
- ponyous 10mo agoJust generated a bunch of 3D CAD models using Gemini 3.0 to see how it compares in spatial understanding and it's heaps better than anything currently out there - not only intelligence but also speed. Will run extended benchmarks later, let me know if you want to see actual data.
- giancarlostoro 10mo agoI'm not familiar enough with CAD what type of format is it?
- ponyous 10mo agoIt’s not a format, but in my mind it implies designs that are supposed to be functional as opposed to models that are meant for virtual games. It generated a blender script that makes the model.
- bilbo0s 10mo agoDid your prompt instruct it to use blender?
- ponyous 10mo agoYes. I’ve been working and refining the prompt for some time now (months). It’s about 10k tokens now.
- JulesRosser 10mo agoWould you mind sharing the prompt please?
- adastra22 10mo agoI would have used OpenSCAD for that purpose.
- 10mo ago
- srameshc 10mo agoI think I am in this AI fatigue phase. I am past all hype with models, tools and agents and back to problem and solution approach, sometimes code gen with AI , sometimes think and ask for a piece of code. But not offloading to AI and buying all the bs, waiting it to do magic with my codebase.
- jstummbillig 10mo agoI think it's fun to see what is not even considered magic anymore today.
- mountainriver 10mo agoPeople would have had a heart attack if they saw this 5 years ago for the first time. Now artificial brains are “meh” :)
- root_axis 10mo agoTrue of almost every new technology.
- abound 10mo agoI hesitate to lump this into the "every new technology" bucket. There are few things that exist today that, similar to what GP said, would have been literal voodoo black magic a few years ago. LLMs are pretty singular in a lot of ways, and you can do powerful things with them that were quite literally impossible a few short years ago. One is free to discount that, but it seems more useful to understand them and their strengths, and use them where appropriate. Even tools like Claude Code have only been fully released for six months, and they've already had a pretty dramatic impact on how many developers work.
- asadotzler 10mo agoMore people got more value out of iPhone, including financially.
- bilekas 10mo ago> The Gemini app surpasses 650 million users per month, more than 70% of our Cloud customers use our AI, 13 million developers have built with our generative models, and that is just a snippet of the impact we’re seeing Not to be a negative nelly, but these numbers are definitely inflated due to Google literally pushing their AI into everything they can, much like M$. Can't even search google without getting an AI response. Surely you can't claim those numbers are legit.
- blinding-streak 10mo agoGemini app != Google search. You're implying they're lying?
- lalitmaganti 10mo ago> Gemini app surpasses 650 million users per month Unless these numbers are just lies, I'm not sure how this is "pushing their AI into everything they can". Especially on iOS where every user is someone who went to App Store and downloaded it. Admittedly on Android, Gemini is preinstalled these days but it's still a choice that users are making to go there rather than being an existing product they happen to user otherwise. Now OTOH "AI overviews now have two billion users" can definitely be criticised in the way you suggest.
- aniforprez 10mo agoI don't know for sure but they have to be counting users like me whose phone has had Gemini force installed on an update and I've only opened the app by accident while trying to figure out how to invoke the old actually useful Assistant app
- bespokedevelopr 10mo agoWow so the polymarket insider bet was true then.. https://old.reddit.com/r/wallstreetbets/comments/1oz6gjp/new_polymarket_account_created_12_days_ago_put/ https://old.reddit.com/r/wallstreetbets/comments/1oz6gjp/new...
- giarc 10mo agoThese prediction markets are so ripe for abuse it's unbelievable. People need to realize there are real people on the other side of these bets. Brian Armstong, CEO of Coinbase intentionally altered the outcome of a bet by randomly stating "Bitcoin, Ethereum, blockchain, staking, Web3" at the end of an earnings call. These types of bets shouldn't be allowed.
- HDThoreaun 10mo agoI’m pretty sure that these model release date markets are made to be abused. They’re just a way to pay insiders to tell you when the model will be released. The mention markets are pure degenerate gambling and everyone involved knows that
- ATMLOTTOBEER 10mo agoCorrect, and this is actually how all markets work in the sense that they allow for price discovery :)
- ethmarks 10mo agoThe point of prediction markets isn't to be fair. They are not the stock market. The point of prediction markets is to predict. They provide a monetary incentive for people who are good at predicting stuff. Whether that's due to luck, analysis, insider knowledge, or the ability to influence the result is irrelevant. If you don't want to participate in an unfair market, don't participate in prediction markets.
- giarc 10mo agoBut what's the point of predicting how many times Elon will say "Trump" on an earnings call (or some random event Kalshi or Polymarket make up)? At least the stock market serves a purpose. People will claim "prediction markets are great for price discovery!" Ok. I'm so glad we found out the chance of Nicki Minaj saying "Bible" during some recent remarks. In case you were wondering, the chance peaked at around 45% and she did not say 'bible'! She passed up a great opportunity to buy the "yes" and make a ton of money! https://kalshi.com/markets/kxminajmention/nicki-minaj/kxminajmention-25nov19 https://kalshi.com/markets/kxminajmention/nicki-minaj/kxmina...
- mikeortman 10mo agoIts available for me now in gemini.google.com.... but its failing so bad at accurate audio transcription. Its transcribing the meeting but hallucinates badly... both in fast and thinking mode. Fast mode only transcribed about a fifth of the meeting before saying its done. Thinking mode completely changed the topic and made up ENTIRE conversations. Gemini 2.5 actually transcribed it decently, just occasional missteps when people talked over each other. I'm concerned.
- coffeecoders 10mo agoFeels like the same consolidation cycle we saw with mobile apps and browsers are playing out here. The winners aren’t necessarily those with the best models, but those who already control the surface where people live their digital lives. Google injects AI Overviews directly into search, X pushes Grok into the feed, Apple wraps "intelligence" into Maps and on-device workflows, and Microsoft is quietly doing the same with Copilot across Windows and Office. Open models and startups can innovate, but the platforms can immediately put their AI in front of billions of users without asking anyone to change behavior (not even typing a new URL).
- Workaccount2 10mo agoAI overviews has arguable done more harm than good for them, because people assume it's Gemini, but really it's some ultra light weight model made for handling millions of queries a minute, and has no shortage of stupid mistakes/hallucinations.
- acoustics 10mo agoMicrosoft hasn't been very quiet about it, at least in my experience. Every time I boot up Windows I get some kind of blurb about an AI feature.
- CobrastanJorji 10mo agoMan, remember the days where we'd lose our minds at our operating systems doing stuff like that?
- esafak 10mo agoThe people who lost their minds jumped ship. And I'm not going to work at a company that makes me use it, either. So, not my problem.
- bitpush 10mo ago> Google injects AI Overviews directly into search, X pushes Grok into the feed, Apple wraps "intelligence" into Maps and on-device workflows, and Microsoft is quietly doing the same with Copilot across Windows and Office. One of them isnt the same as others (hint: It is Apple). The only thing Apple is doing with Maps is, is adding ads https://www.macrumors.com/2025/10/26/apple-moving-ahead-with-ads-in-maps/ https://www.macrumors.com/2025/10/26/apple-moving-ahead-with...
- mccoyb 10mo agoI truly do not understand what plan to use so I can use this model for longer than ~2 minutes. Using Anthropic or OpenAI's models are incredibly straightforward -- pay us per month, here's the button you press, great. Where do I go for this for these Google models?
- kachapopopow 10mo agoai studio, you get a bunch of usage free if you want more you buy credits (google one subscriptions also give you some additional usage)
- mccoyb 10mo agoI see -- so this is the "paid" AI studio plan? Does that have any relation to the Gemini plan thing: https://one.google.com/explore-plan/gemini-advanced?utm_source=gemini&utm_medium=web&utm_campaign=sidenav_evo&g1_landing_page=65 https://one.google.com/explore-plan/gemini-advanced?utm_sour... ?
- kachapopopow 10mo agothat's for the first party google integrations - not 3rd party. ai studio just gives you an api key that you can use anywhere.
- fschuett 10mo agoUpdate VSCode to the latest version and click the small "Chat" button at the top bar. GitHub gives you like $20 for free per month and I think they have a deal with the larger vendors because their pricing is insanely cheap. One week of vibe-coding costs me like $15, only downside to Copilot is that you can't work on multiple projects at the same time because of rate-limiting.
- mccoyb 10mo agoI'm asking about Gemini, not Copilot.
- deanc 10mo agoThe AntiGravity seems to be a bit overwhelmed. Unable to set up an account at the moment.
- NullCascade 10mo agoI'm not a mathematician but I think we underestimate how useful pure mathematics can be to tell whether we are approaching AGI. Can the mathematicians here try ask it to invent new novel math related to [Insert your field of specialization] and see if it comes up with something new and useful? Try lowering the temperature, use SymPy etc.
- ducttapecrown 10mo agoTerry Tao is writing about this on his blog.
- stevesimmons 10mo agoA nice Easter egg in the Gemini 3 docs [1]: If you are transferring a conversation trace from another model, ... to bypass strict validation in these specific scenarios, populate the field with this specific dummy string: "thoughtSignature": "context_engineering_is_the_way_to_go" [1] https://ai.google.dev/gemini-api/docs/gemini-3?thinking=high#migrating_from_other_models https://ai.google.dev/gemini-api/docs/gemini-3?thinking=high...
- bijant 10mo agoIt's an artifact of the problem that they don't show you the reasoning output but need it for further messages so they save each api conversation on their side and give you a reference number. It sucks from a GDPR compliance perspective as well as in terms of transparent pricing as you have no way to control reasoning trace length (which is billed at the much higher output rate) other than switching between low/high but if the model decides to think longer "low" could result in more tokens used than "high" for a prompt where the model decides not to think that much. "thinking budgets" are now "legacy" and thus while you can constrain output length you cannot constrain cost. Obviously you also cannot optimize your prompts if some red herring makes the LLM get hung up on something irrelevant only to realize this in later thinking steps. This will happen with EVERY SINGLE prompt if it's caused by something in your system prompt. Finding what makes the model go astray can be rather difficult with 15k token system prompts or a multitude of MCP tools, you're basically blinded while trying to optimize a black box. Obviously you can try different variations of different parts of your system prompt or tool descriptions but just because they result in less thinking tokens does not mean they are better if those reasoning steps where actually beneficial (if only in edge cases) this would be immediately apparent upon inspection but hard/impossible to find out without access to the full Chain of Thought. For the uninitiated, the reasons OpenAI started replacing the CoT with summaries, were A. to prevent rapid distillation as they suspected deepSeek to have used for R1 and B. to prevent embarrassment if App users see the CoT and find parts of it objectionable/irrelevant/absurd (reasoning steps that make sense for an LLM do not necessarily look like human reasoning). That's a tradeoff that is great with end-users but terrible for developers. As Open Weights LLMs necessarily output their full reasoning traces the potential to optimize prompts for specific tasks is much greater and will for certain applications certainly outweigh the performance delta to Google/OpenAI.
- scrollop 10mo agoHere it makes a text based video editor that works: https://youtu.be/MPjOQIQO8eQ?si=wcrCSLYx3LjeYDfi&t=797 https://youtu.be/MPjOQIQO8eQ?si=wcrCSLYx3LjeYDfi&t=797
- mpeg 10mo agoWell, it just found a bug in one shot that Gemini 2.5 and GPT5 failed to find in relatively long sessions. Claude 4.5 had found it but not one shot. Very subjective benchmark, but it feels like the new SOTA for hard tasks (at least for the next 5 minutes until someone else releases a new model)
- jordanpg 10mo agoWhat is Gemini 3 under the hood? Is it still just a basic LLM based on transformers? Or are there all kinds of other ML technologies bolted on now? I feel like I've lost the plot.
- meowface 10mo agoI am very ignorant in this field but I am pretty sure under the hood they are all still fundamentally built on the transformer architecture, or at least innovations on the original transformer architecture.
- anilgulecha 10mo agoIt's a mixture-of-experts model. Basically N smaller model pieces put together, and when inference occurs, only 1 is active at a time. Each model piece would be tuned/good in one area.
- becquerel 10mo agoThe industry is still seeing how far they can take transformers. We've yet to reach a dollar value where it stops being worth pumping money into them.
- tylervigen 10mo agoI am personally impressed by the continued improvement in ARC-AGI-2, where Gemini 3 got 31.1% (vs ChatGPT 5.1's 17.6%). To me this is the kind of problem that does not lend itself well to LLMs - many of the puzzles test the kind of thing that humans intuit because of millions of years of evolution, but these concepts do not necessarily appear in written form (or when they do, it's not clear how they connect to specific ARC puzzles). The fact that these models can keep getting better at this task given the setup of training is mind-boggling to me. The ARC puzzles in question: https://arcprize.org/arc-agi/2/ https://arcprize.org/arc-agi/2/
- grantpitt 10mo agoAgreed, it also leads performance on arc-agi-1. Here's the leaderboard where you can toggle between arc-agi-1 and 2: https://arcprize.org/leaderboard https://arcprize.org/leaderboard
- energy123 10mo agoIt leads on arc-agi-1 with Gemini 3.0 Deep Think, which uses "tool calls" according to google's post, whereas regular Gemini 3.0 Pro doesn't use "tool calls" for the same benchmark. I am unsure how significant this difference is.
- stephc_int13 10mo agoWhat I would do if I was in the position of a large company in this space is to arrange an internal team to create an ARC replica, covering very similar puzzles and use that as part of the training. Ultimately, most benchmarks can be gamed and their real utility is thus short-lived. But I think this is also fair to use any means to beat it.
- tylervigen 10mo agoI agree that for any given test, you could build a specific pipeline to optimize for that test. I supposed that's why it is helpful to have many tests. However, many people have worked hard to optimize tools specifically for ARC over many years, and it's proven to be a particularly hard test to optimize for. This is why I find it so interesting that LLMs can do it well at all, regardless of whether tests like it are included in training.
- dankobgd 10mo agoevery day, new game changer
- deleted 10mo ago[deleted]
- hubraumhugo 10mo agoNo gemini-3-flash yet, right? Any ETA on that mentioned? 2.5-flash has been amazing in terms of cost/value ratio.
- 8note 10mo agoive found gemini 2.5-flash works better (for.agentic coding) than pro, too
- casey2 10mo agoThe first paragraph is pure delusion. Why do investors like delusional CEOs so much? I would take it as a major red flag.
- clusterhacks 10mo agoI wish I could just pay for the model and self-host on local/rented hardware. I'm incredibly suspicious of companies totally trying to capture us with these tools.
- lfx 10mo agoTechnically you can! I haven't seen it in the box yet, and pricing is unknown https://cloud.google.com/blog/products/ai-machine-learning/run-gemini-and-ai-on-prem-with-google-distributed-cloud https://cloud.google.com/blog/products/ai-machine-learning/r...
- clusterhacks 10mo agoThat's interesting. While I suspect the pricing will lean heavily into enterprise sales rather than personal licenses, I personally like the idea buying models that I then own and control. Any steps from companies that make that more possible is great.
- slackerIII 10mo agoWhat's the easiest way to set up automatic code review for PRs for my team on GitHub using this model?
- esafak 10mo agohttps://github.com/marketplace/gemini-code-assist https://github.com/marketplace/gemini-code-assist
- colechristensen 10mo agoAsk it. If it's good enough to be useful on your code base, it better be good enough to instruct you on how to use it. How easy it is depends on whether or not they've built that kind of thing in
- qustrolabe 10mo agoOut of all other companies Google provide the most generous free access so far. I bet this gives them plenty of data to train even better models
- serjester 10mo agoIt's disappointing there's no flash / lite version - this is where Google has excelled up to this point.
- aoeusnth1 10mo agoMaybe they're slow rolling the announcements to be in the news more
- coffeebeqn 10mo agoMost likely. And/or they use the full model to train the smaller ones somehow
- FergusArgyll 10mo agoThe term of art is distillation
- icapybara 10mo agoAnyone know how Gemini CLI with this model compares to Codex and Claude Code?
- dwringer 10mo agoWell, I tried a variation of a prompt I was messing with in Flash 2.5 the other day in a thread about AI-coded analog clock faces. Gemini Pro 3 Preview gave me a result far beyond what I saw with Flash 2.5, and got it right in a single shot.[0] I can't say I'm not impressed, even though it's a pretty constrained example. > Please generate an analog clock widget, synchronized to actual system time, with hands that update in real time and a second hand that ticks at least once per second. Make sure all the hour markings are visible and put some effort into making a modern, stylish clock face. Please pay attention to the correct alignment of the numbers, hour markings, and hands on the face. [0] https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%5B%221W5xcsbljPVcaYTveFTdvOWdiqh_PgOF2%22%5D,%22action%22:%22open%22,%22userId%22:%22105800868059822502362%22,%22resourceKeys%22:%7B%7D%7D&usp=sharing https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%...
- thegrim33 10mo ago"Allow access to Google Drive to load this Prompt." .... why? For what possible reason? No, I'm not going to give access to my privately stored file share in order to view a prompt someone has shared. Come on, Google.
- LiamPowell 10mo agoYou don't want to give Google access to files you've stored in Google Drive? It's also only access to an application specific folder, not all files.
- tibbar 10mo agoWell, you also have to allow it to train on your data. Although this is not explicitly about your Google drive data, and probably requires you to submit a prompt yourself, the barriers here are way to weak/fuzzy for me consider granting access via any account with private info.
- lxgr 10mo agoBecause most likely (at least according to Hanlon's razor) they somehow decided that using Google Drive as the only persistent storage backing AI studio was a reasonable UX decision. It probably makes some sense internally in big tech corporation logic (no new data storage agreements on top of the ones the user has already agreed to when signing up for Drive etc.), but as a user, I find it incredibly strange too – especially since the text chats are in some proprietary format I can't easily open on my local GDrive replica, but the images generated or uploaded just look like regular JPEGs and PNGs.
- siva7 10mo agoI have my own private benchmarks for reasoning capabilities on complex problems and i test them against SOTA models regularly (professional cases from law and medicine). Anthropic (Sonnet 4.5 Extended Thinking) and OpenAI (Pro Models) get halfway decent results on many cases while Gemini Pro 2.5 struggled (it was overconfident in its initial assumptions). So i ran these benchmarks against Gemini 3 Pro and i'm not impressed. The reasoning is way more nuanced than their older model but it still makes mistakes which the other two SOTA competitor models don't make. Like it forgets in a law benchmark that those principles don't apply in the country from the provided case. It seems very US centric in its thinking whereas Anthropic and OpenAI pro models seem to be more aware around the context of assumed culture from the case. All in - i don't think this new model is ahead of the other two main competitors - but it has a new nuanced touch and is certainly way better than Gemini 2.5 pro (which is more telling how bad actually that one was for complex problems).
- MaxL93 10mo ago> It seems very US centric in its thinking I'm not surprised. I'm French and one thing I've consistently seen with Gemini is that it loves to use Title Case (Everything is Capitalized Except the Prepositions) even in French or other languages where there is no such thing. A 100% american thing getting applied to other languages by the sheer power of statistical correlation (and probably being overtrained on USA-centric data). At the very least it makes it easy to tell when someone is just copypasting LLM output into some other website.
- mpalmer 10mo ago> Title Case (Everything is Capitalized Except the Prepositions) If this is an American thing I'm happy to disown/denounce it; it's my least favorite pattern in Gemini output.
- irthomasthomas 10mo agoI asked it to summarize an article about the Zizians which mentions Yudkowsky SEVEN times. Gemini-3 did not mention him once. Tried it ten times and got zero mention of Yudkowsky, despite him being a central figure in the story. https://xcancel.com/xundecidability/status/1990828697088131145 https://xcancel.com/xundecidability/status/19908286970881311... Also, can you guess which pelican SVG was gemini 3 vs 2.5? https://xcancel.com/xundecidability/status/1990811319172321300 https://xcancel.com/xundecidability/status/19908113191723213...
- briga 10mo agoMaybe it has guard rails against such things? That would be my main guess on the Zizian one.
- gregsadetsky 10mo agoInteresting, yeah! Just tried "summarize this story and list the important figures from it" with Gemini 2.5 Pro and 3 and they both listed 10 names each, but without including Yudkowsky. Asking the follow up "what are ALL the individuals mentioned in the story" results in both models listing ~40 names and both of those lists include Yudkowsky.
- stickfigure 10mo agoHe's not a central figure in the narrative, he's a background character. Things he created (MIRI, CFAR, LessWrong) are important to the narrative, the founder isn't. If I had to condense the article, I'd probably cut him out too. Summarization is inherently lossy.
- irthomasthomas 10mo ago> Eliezer Yudkowsky is a central figure in the article, mentioned multiple times as the intellectual originator of the community from which the "Zizians" splintered. His ideas and organizations are foundational to the entire narrative.
- stickfigure 10mo ago
- alach11 10mo agoThis is a really impressive release. It's probably the biggest lead we've seen from a model since the release of GPT-4. Seems likely that OpenAI rushed out GPT-5.1 to beat the Gemini 3 release, knowing that their model would underperform it.
- m3kw9 10mo agoIf it ain't quantum leap, new models are just "OS updates".
- sunaookami 10mo agoGemini CLI crashes due to this bug: https://github.com/google-gemini/gemini-cli/issues/13050 https://github.com/google-gemini/gemini-cli/issues/13050 and when applying the fix in the settings file I can't login with my Google account due to "The authentication did not complete successfully. The following products are not yet authorized to access your account" with useless links to completely different products (Code Assist). Antigravity uses Open-VSX and can't be configured differently even though it says it right there (setting is missing). Gemini website still only lists 2.5 Pro. Guess I will just stick to Claude.
- bityard 10mo ago> Whether you’re an experienced developer or a vibe coder I absolutely LOVE that Google themselves drew a sharp distinction here.
- rafaquintanilha 10mo agoYou realize this is copy to attract more people to the product, right?
- jstummbillig 10mo agoHow could they.
- WXLCKNO 10mo agoValve could learn from Google here
- deleted 10mo ago[deleted]
- pflenker 10mo ago> Since then, it’s been incredible to see how much people love it. AI Overviews now have 2 billion users every month. Come on, you can’t be serious.
- muzani 10mo agoThis is so disingenuous that it hurts the credibility of the whole thing.
- XCSme 10mo agoHow's the pelican?
- briga 10mo agoEvery big new model release we see benchmarks like ARC and Humanity's Last Exam climbing higher and higher. My question is, how do we know that these benchmarks are not a part of the training set used for these models? It could easily have been trained to memorize the answers. Even if the datasets haven't been copy pasted directly, I'm sure it has leaked onto the internet to some extent. But I am looking forward to trying it out. I find Gemini to be great as handling large-context tasks, and Google's inference costs seem to be among the cheapest.
- stephc_int13 10mo agoEven if the benchmark themselves are kept secret, the process to create them is not that difficult and anyone with a small team of engineers could make a replica in their own labs to train their models on. Given the nature of how those models work, you don't need exact replicas.
- poemxo 10mo agoIt's amazing to see Google take the lead while OpenAI worsens their product every release.
- vivzkestrel 10mo agohas anyone managed to use any of the AI models to build a complete 3D fps game using web GL or open GL?
- kridsdale3 10mo agoI made a webgl copy of wolfenstein with prompt engineering in browser-based "Make a website" tool that was gemini-powered.
- vivzkestrel 10mo agomind sharing what tool that was that lets you run gemini on the browser in interactive mode to make games?
- lairv 10mo agoOut of curiosity, I gave it the latest project euler problem published on 11/16/2025, very likely out of the training data Gemini thought for 5m10s before giving me a python snippet that produced the correct answer. The leaderboard says that the 3 fastest human to solve this problem took 14min, 20min and 1h14min respectively Even thought I expect this sort of problem to very much be in the distribution of what the model has been RL-tuned to do, it's wild that frontier model can now solve in minutes what would take me days
- orly01 10mo agoWow. Sounds pretty impressive.
- qsort 10mo agoTo be fair a lot of the impressive Elo scores models get are simply due to the fact that they're faster: many serious competitive coders could get the same or better results given enough time. But seeing these results I'd be surprised if by the end of the decade we don't have something that is to these puzzles what Stockfish is to chess. Effectively ground truth and often coming up with solutions that would be absolutely ridiculous for a human to find within a reasonable time limit.
- nerdsniper 10mo agoI’d love if anyone could provide examples of such AND(“ground truth”, “absolutely ridiculous”) solutions! Even if they took clever humans a long time to create. I’m curious to explore such fun programming code. But I’m also curious to explore what knowledgeable humans consider to be both “ground truth” as well as “absolutely ridiculous” to create within the usual time constraints.
- qsort 10mo agoI'm not explaining myself right. Stockfish is a superhuman chess program. It's routinely used in chess analysis as "ground truth": if Stockfish says you've made a mistake, it's almost certain you did in fact make a mistake[0]. Also, because it's incomparably stronger than even the very best humans, sometimes the moves it suggests are extremely counterintuitive and it would be unrealistic to expect a human to find them in tournament conditions. Obviously software development in general is way more open-ended, but if we restrict ourselves to puzzles and competitions, which are closed game-like environments, it seems plausible to me that a similar skill level could be achieved with an agent system that's RL'd to death on that task. If you have base models that can get there, even inconsistently so, and an environment where making a lot of attempts is cheap, that's the kind of setup that RL can optimize to the moon and beyond. I don't predict the future and I'm very skeptical of anybody who claims to do so, correctly predicting the present is already hard enough, I'm just saying that given the progress we've already made I would find plausible that a system like that could be made in a few years. The details of what it would look like are beyond my pay grade. --- [0] With caveats in endgames, closed positions and whatnot, I'm using it as an example.
- yomismoaqui 10mo agoFrom an initial testing of my personal benchmark it works better than Gemini 2.5 pro. My use case is using Gemini to help me test a card game I'm developing. The model simulates the board state and when the player has to do something it asks me what card to play, discard... etc. The game is similar to something like Magic the Gathering or Slay the Spire with card play inspired by Marvel Champions (you discard cards from your hand to pay the cost of a card and play it) The test is just feeding the model the game rules document (markdown) with a prompt asking it to simulate the game delegating the player decisions to me, nothing special here. It seems like it forgets rules less than Gemini 2.5 Pro using thinking budget to max. It's not perfect but it helps a lot to test little changes to the game, rewind to a previous turn changing a card on the fly, etc...
- pgroves 10mo agoI was hoping Bash would go away or get replaced at some point. It's starting to look like it's going to be another 20 years of Bash but with AI doodads.
- __MatrixMan__ 10mo agoNushell scratches the itch for me 95% of the time. I haven't yet convinced anybody else to make the switch, but I'm trying. Haven't yet fixed the most problematic bug for my useage, but I'm trying. What are you doing to help kill bash?
- energy123 10mo agoImpressive. Although the Deep Think benchmark results are suspicious given they're comparing apples (tools on) with oranges (tools off) in their chart to visually show an improvement.
- syedshahmir7214 10mo agoI think from last few releases of these models from all companies, I have not observed much improvements in the response of these models. Their claims and launches are a little over hyped.
- nextworddev 10mo agoIt’s over for Anthropic. That’s why Google’s cool with Claude being on Azure. Also probably over for OpenAI
- nighwatch 10mo agoI just tested the Gemini 3 preview as well, and its capabilities are honestly surprising. As an experiment I asked it to recreate a small slice of Zelda , nothing fancy, just a mock interface and a very rough combat scene. It managed to put together a pretty convincing UI using only SVG, and even wired up some simple interactions. It’s obviously nowhere near a real game, but the fact that it can structure and render something that coherent from a single prompt is kind of wild. Curious to see how far this generation can actually go once the tooling matures.
- crawshaw 10mo agoHas anyone who is a regular Opus / GPT5-Codex-High / GPT5 Pro user given this model a workout? Each Google release is accompanied by a lot of devrel marketing that sounds impressive but whenever I put the hours into eval myself it comes up lacking. Would love to hear that it replaces another frontier model for someone who is not already bought into the Gemini ecosystem.
- film42 10mo agoAt this point I'm only using google models via Vertex AI for my apps. They have a weird QoS rate limit but in general Gemini has been consistently top tier for everything I've thrown at it. Anecdotal, but I've also not experienced any regression in Gemini quality where Claude/OpenAI might push iterative updates (or quantized variants for performance) that cause my test bench to fail more often.
- gordonhart 10mo agoMatches my experience exactly. It's not the best at writing code but Gemini 2.5 Pro is (was) the hands-down winner in every other use case I have. This was hard for me to accept initially as I've learned to be anti-Google over the years, but the better accuracy was too good to pass up on. Still expecting a rugpull eventually — price hike, killing features without warning, changing internal details that break everything — but it hasn't happened yet.
- Szpadel 10mo agoI gave it a spin with instructions that worked great with gpt-5-codex (5.1 regressed a lot so I do not even compare to it). Code quality was fine for my very limited tests but I was disappointed with instruction following. I tried few tricks but I wasn't able to convince it to first present plan before starting implementation. I have instructions describing that it should first do exploration (where it tried to discover what I want) then plan implementation and then code, but it always jumps directly to code. this is bug issue for me especially because gemini-cli lacks plan mode like Claude code. for codex those instructions make plan mode redundant.
- eknkc 10mo agoLooks like it is already available on VSCode Copilot. Just tried a prompt that was not returning anything good on Sonnet 4.5. (Did not spend much time though, but the prompth was already there on the chat screen so I switched the model and sent it again) Gemini 3 worked much better and I actually committed the changes that it created. I don't mean its revolutionary or anything but it provided a nice summary of my request and created a decent simple solution. Sonnet had created a bunch of overarching changes that I would not even bother reviewing. Seems nice. Will probably use it for 2 weeks until someone else releases a 1.0001x better model.
- flyinglizard 10mo agoYou were probably stuck at some local model minima avoidable by simply changing the model to something else.
- zone411 10mo agoSets a new record on the Extended NYT Connections benchmark: 96.8 (https://github.com/lechmazur/nyt-connections/ https://github.com/lechmazur/nyt-connections/). Grok 4 is at 92.1, GPT-5 Pro at 83.9, Claude Opus 4.1 Thinking 16K at 58.8. Gemini 2.5 Pro scored 57.6, so this is a huge improvement.
- dudeinhawaii 10mo agoGemini has been so far behind agentically it's comical. I'll be giving it a shot but it has a herculean task ahead of itself. It has to not only be "good enough" but a "quantum leap forward". That said, OpenAI was in the same place earlier in the year and very quickly became the top agentic platform with GPT-5-Codex. The AI crowd is surprisingly not sticky. Coders quickly move to whatever the best model is. Excited to see Gemini making a leap here.
- ryandrake 10mo agoI don't even know what the fuck "agentic" is or why the hell I would want it all over my software. So tired of everything in the computing world today.
- esafak 10mo agoPrompting, planning, iteration, coding, and tool use over an entire code base until a problem is solved.
- rkozik1989 10mo agoSounds like an antipattern being rebranded as a solution. I shouldn't have to precisely instruct AI on how to solve every problem. I should be able to give it requirements and with its vast knowledge it should be able to understand various design elements within a system like design patterns and make the appropriate change without me needing to tell it to look for those things.
- ur-whale 10mo ago> So tired of everything in the computing world today. That's actually sad, and if you're - like I am - long in the tooth in computer land, you should definitely try agentic in CLI mode. I haven't been that excited to play with a computer in 30 years.
- SchemaLoad 10mo agoAs far as I can tell, it just means giving the LLM the ability to run commands, read files, edit files, and run in a loop until some goal is achieved. Compared to chat interfaces where you just input text and get one response back.
- aerhardt 10mo agoCombining structured outputs with search is the API feature I was looking for. Honestly crazy that it wasn’t there to start with - I have a project that is mostly Gemini API but I’ve had to mix in GPT-5 just for this feature. I still use ChatGPT and Codex as a user but in the API project I’ve been working on Gemini 2.5 Pro absolutely crushed GPT-5 in the accuracy benchmarks I ran. As it stands Gemini is my de facto standard for API work and I’ll be following very closely the performance of 3.0 in coming weeks.
- syspec 10mo agoI have "unlimited" access to both Gemini 2.5 Pro and Claude 4.5 Sonnet through work. From my experience, both are capable and can solve nearly all the same complex programming requests, but time and time again Gemini spits out reams and reams of code so over engineered, that totally works, but I would never want to have to interact with. When looking at the code, you can't tell why it looks "gross", but then you ask Claude to do the same task in the same repo (I use Cline, it's just a dropdown change) and the code also works, but there's a lot less of it and it has a more "elegant" feeling to it. I know that isn't easy to capture in benchmarks, but I hope Gemini 3.0 has improved in this regard
- jmkni 10mo agoI can relate to this, it's doing exactly what I want, but it ain't pretty. It's fine though if you take the time to learn what it's doing and write a nicer version of it yourself
- poyu 10mo agobut I would never want to have to interact with That is its job security ;)
- eitally 10mo agoI have had a similar experience vibe coding with Copilot (ChatGPT) in VSCode, against the Gemini API. I wanted to create a dad joke generator and then have it also create a comic styled 4 cel interpretation of the joke. Simple, right? I was able to easily get it to create the joke, but it repeatedly failed on the API call for the image generation. What started as perhaps 100 lines of total code in two files ended up being about 1500 LOC with an enormous built-in self-testing mechanism ... and it still didn't work.
- plaidfuji 10mo agoI have the same experience with Gemini, that it’s incredibly accurate but puts in defensive code and error handling to a fault. It’s pretty easy to just tell it “go easy on the defensive code” / “give me the punchy version” and it cleans it up
- catigula 10mo agoThe problem with experiencing LLM releases nowadays is that it is no longer trivial to understand the differences in their vast intelligences so it takes awhile to really get a handle on what's even going on.
- SXX 10mo agoStatic Pelican is boring. First attempt: Generate SVG animation of following: 1 - There is High fantasy mage tower with a top window a dome 2 - Green goblin come in front of tower with a torch 3 - Grumpy old mage with beard appear in a tower window in high purple hat 4 - Mage sends fireball that burns goblin and all screen is covered in fire. Camera view must be from behind of goblin back so we basically look at tower in front of us: https://codepen.io/Runway/pen/WbwOXRO https://codepen.io/Runway/pen/WbwOXRO
- Rudybega 10mo agoHoly crap. That's actually kind of incredible for a first attempt.
- sosodev 10mo agoWow, that's very impressive
- SXX 10mo agoAfter few more attempts longer animation with a story from my gamedev inspired mind: https://codepen.io/Runway/pen/zxqzPyQ https://codepen.io/Runway/pen/zxqzPyQ PS: but yeah thats attempt #20 or something.
- fatty_patty89 10mo agoSeizure warning for the above link edit: flashing lights at the end seem to be mostly becauseo f darkreader extension
- fromwilliam 10mo agoThis is honestly incredible
- nyantaro1 10mo agowe are so cooked
- WithinReason 10mo ago
- John-Tony 10mo ago[flagged]
- jennyholzer 10mo agoboooooooooooooo
- jennyholzer 10mo ago"AI" benchmarks are and have consistently been lies and misinformation. Gemini is dead in the water.
- ilaksh 10mo agookay since Gemini 3 is AI mode now, I switched from the free perplexity back to google as being my search default.
- DanMcInerney 10mo agoA 50% increase over ChatGPT 5.1 on ARC-AGI2 is astonishing. If that's true and representative (a big if), it lends credence to this being the first of the very consistent agentically-inclined models because it's able to follow a deep tree of reasoning to solve problems accurately. I've been building agents for a while and thus far have had to add many many explicit instructions and hardcoded functions to help guide the agents in how to complete simple tasks to achieve 85-90% consistency.
- puttycat 10mo agoWhere is this figure taken from?
- machiaweliczny 10mo agoI think it's due to improvements in vision basically, the arc agi 2 is very visual
- machiaweliczny 10mo agoVision is very far from solved IMO, simple modifications to inputs results in high differences still, lines aren't recognized etc..
- zone411 10mo agoSets a new record on the Extended NYT Connections: 96.8. Gemini 2.5 Pro scored only 57.6. https://github.com/lechmazur/nyt-connections/ https://github.com/lechmazur/nyt-connections/
- CephalopodMD 10mo agoWhat I'm getting from this thread is that people have their own private benchmarks. It's almost a cottage industry. Maybe someone should crowd source those benchmarks, keep them completely secret, and create a new public benchmark of people's private AGI tests. All they should release for a given model is the final average score.
- JacobiX 10mo agoTested it on a bug that Claude and ChatGPT Pro struggled with, it nailed it, but only solved it partially (it was about matching data using a bipartite graph). Another task was optimizing a complex SQL script: the deep-thinking mode provided a genuinely nuanced approach using indexes and rewriting parts of the query. ChatGPT Pro had identified more or less the same issues. For frontend development, I think it’s obvious that it’s more powerful than Claude Code, at least in my tests, the UIs it produces are just better. For backend development, it’s good, but I noticed that in Java specifically, it often outputs code that doesn’t compile on the first try, unlike Claude.
- AstroBen 10mo agoFirst impression is I'm having a distinctly harder time getting this to stick to instructions as compared to Gemini 2.5
- mrinterweb 10mo agoHit the Gemini 3 quota on the second prompt in antigravity even though I'm a pro user. I highly doubt I hit a context window based on my prompt. Hopefully, it is just first day of near general availability jitters.
- deleted 10mo ago[deleted]
- BoorishBears 10mo agoSo they won't release multimodal or Flash at launch, but I'm guessing people who blew smoke up the right person's backside on X are already building with it Glad to see Google still can't get out of its own way.
- BoorishBears 10mo agoI don't want to be one of those assholes who only calls out when they were right: I was very wrong
- John-Tony12 10mo ago[flagged]
- oceanplexian 10mo agoSuspicious that none of the benchmarks include Chinese models even they scored higher on the benchmarks than the models they are comparing to?
- recitedropper 10mo agoWho wants to bet they benchmaxxed ARC-AGI-2? Nothing in their release implies they found some sort of "secret sauce" that justifies the jump. Maybe they are keeping that itself secret, but more likely they probably just have had humans generate an enormous number of examples, and then synthetically build on that. No benchmark is safe, when this much money is on the line.
- HarHarVeryFunny 10mo agoI'd also be curious what kind of tools they are providing to get the jump from Pro to Deep Think (with tools) performance. ARC-AGI specialized tools?
- horhay 10mo agoThey ran the tests themselves only on semi-private evals. Basically the same caveat as when o3 supposedly beat ARC1
- sosodev 10mo agoHere's some insight from Jeff Dean and Noam Shazeer's interview with Dwarkesh Patel https://youtu.be/v0gjI__RyCY&t=7390 https://youtu.be/v0gjI__RyCY&t=7390 > When you think about divulging this information that has been helpful to your competitors, in retrospect is it like, "Yeah, we'd still do it," or would you be like, "Ah, we didn't realize how big a deal transformer was. We should have kept it indoors." How do you think about that? > Some things we think are super critical we might not publish. Some things we think are really interesting but important for improving our products; We'll get them out into our products and then make a decision.
- recitedropper 10mo agoI'm sure each of the frontier labs have some secret methods, especially in training the models and the engineering of optimizing inference. That said, I don't think them saying they'd keep a big breakthrough secret would be evidence in this case of a "secret sauce" on ARC-AGI-2. If they had found something fundamentally new, I doubt they would've snuck it into Gemini 3. Probably would cook on it longer and release something truly mindblowing. Or, you know, just take over the world with their new omniscient ASI :)
- testfrequency 10mo agoI continue to not use Gemini as I can’t have my data not trained but also have chat history at the same time. Yes, I know the Workspaces workaround, but that’s silly.
- alksdjf89243 10mo agoPretty obvious how contaminated this site is with goog employees upvoting nonsense like this.
- _2d30 10mo agoGemini 3 is crushing my personal evals for research purposes. I would cancel my ChatGPT sub immediately if Gemini had a desktop app and may still do so if it continues to impress my as much as it has so far and I will live without the desktop app. It's really, really, really good so far. Wow. Note that I haven't tried it for coding yet!
- ethmarks 10mo agoGenuinely curious here: why is the desktop app so important? I completely understand the appeal of having local and offline applications, but the ChatGPT desktop app doesn't work without an internet connection anyways. Is it just the convenience? Why is a dedicated desktop app so much better than just opening a browser tab or even using a PWA? Also, have you looked into open-webui or Msty or other provider-agnostic LLM desktop apps? I personally use Msty with Gemini 2.5 Pro for complex tasks and Cerebras GLM 4.6 for fast tasks.
- _2d30 10mo agoI have a few reasons for the preference: (1) The ability to add context via a local apps integration into OS level resources is big. With Claude, eg, I hit Option-SPC which brings up a prompt bar. From there, taking a screenshot that will get sent my prompt is as simple as dragging a bounding box. This is great. Beyond that, I can add my own MCP connectors and give my desktop app direct access to relevant context in a way that doesn't work via web UI. It may also be inconvenient to give context to a web UI in some case where, eg, I may have a folder of PDFs I want it to be able to reference. (2) Its own icon that I can CMD-TAB to is so much nicer. Maybe that works with a PWA? Not really sure. (3) Even if I can't use an LLM when offline, having access to my chats for context has been repeatedly valuable to me. I haven't looked at provider-agnostic apps and, TBH, would be wary of them.
- ethmarks 10mo ago> The ability to add context via a local apps integration into OS level resources is big Good point. I can see why integrated support for local filesystem tools would be useful, even though I prefer manually uploading specific files to avoid polluting the context with irrelevant info. > Its own icon that I can CMD-TAB to is so much nicer Fair enough. I personally prefer Firefox's tab organization to my OS's window organization, but I can see how separating the LLM into its own window would be helpful. > having access to my chats for context has been repeatedly valuable to me. I didn't at all consider this. Point ceded. > I haven't looked at provider-agnostic apps and, TBH, would be wary of them. Interesting. Why? Is it security? The ones I've listed are open source and auditable. I'm confident that they won't steal my API keys. Msty has a lot of advanced functionality that I haven't seen in other interfaces like allowing you to compare responses between different LLMs, export the entire conversation to Markdown, and edit the LLM's response to manage context. It also sidesteps the problem of '[provider] doesn't have a desktop app' because you can use any provider API.
- realty_geek 10mo agoI would like to try controlling my browser with this model. Any ideas how to do this. Ideally I would like something like openAI's atlas or perplexity's comet but powered by gemini 3.
- ZeroCool2u 10mo agoSeems like their new Antigravity IDE specifically has this built in. https://antigravity.google/docs/browser https://antigravity.google/docs/browser
- realty_geek 10mo agoWow, that is awesome.
- xnx 10mo agoGemini CLI can also control a browser: https://github.com/ChromeDevTools/chrome-devtools-mcp https://github.com/ChromeDevTools/chrome-devtools-mcp
- I_am_tiberius 10mo agoI still need a google account to use it and it always asks me for a phone verification, which I don't want to give to google. That prevents me from using Gemini. I would even pay for it.
- gpm 10mo ago> I would even pay for it. Is it just me or is it generally the case that to pay for anything on the internet you have to enter credit card information including a phone number.
- I_am_tiberius 10mo agoYou never have to add your phone number in order to pay.
- gpm 10mo agoWhile I haven't tried leaving the field blank on every credit card form I've come across, I'm certain that at least some of them considered it required. Perhaps its country specific?
- I_am_tiberius 10mo agoI've never been asked a phone number. Maybe country specific. no idea.
- senfiaj 10mo agoHaven't used Gemini much, but when I used, it often refused to do certain things that ChatGPT did happily. Probably because it has many things heavily censored. Obviously, a huge company like Google is under much heavier regulations than ChatGPT. Unfortunately this greatly reduces its usefulness in many situations despite that Google has more resources and computational power than OpenAI.
- simonw 10mo agoHere are my notes and pelican benchmark, including a new, harder benchmark because the old one was getting too easy: https://simonwillison.net/2025/Nov/18/gemini-3/ https://simonwillison.net/2025/Nov/18/gemini-3/
- libraryofbabel 10mo agoI was interested (and slightly disappointed) to read that the knowledge cutoff for Gemini 3 is the same as for Gemini 2.5: January 2025. I wonder why they didn't train it on more recent data. Is it possible they use the same base pre-trained model and just fine-tuned and RL-ed it better (which, of course, is where all the secret sauce training magic is these days anyhow)? That would be odd, especially for a major version bump, but it's sort of what having the same training cutoff points to?
- simonw 10mo agoThe model card says: https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf https://storage.googleapis.com/deepmind-media/Model-Cards/Ge... > This model is not a modification or a fine-tune of a prior model. I'm curious why they decided not to update the training data cutoff date too.
- stocksinsmocks 10mo agoMaybe that date is a rule of thumb for when AI generated content became so widespread that it is likely to have contaminated future data. Given that people have spoofed authentic Reddit users with Markov chains, it probably doesn’t go back nearly far enough.
- skylurk 10mo agoThey've been training for months to draw that pelican, just for you to move the goalposts.
- oblio 10mo ago
- Retr0id 10mo ago> it’s been incredible to see how much people love it. AI Overviews now have 2 billion users every month Do regular users know how to disable AI Overviews, if they don't love them?
- Dquiroga 10mo agoI asked Gemini to write "a comment response to this thread. I want to start an intense discussion". Gemini 3: The cognitive dissonance in this thread is staggering. We are sitting here cheering for a model that effectively closes the loop on Google’s total information dominance, while simultaneously training our own replacements. Two things in this thread should be terrifying, yet are being glossed over in favor of "look at the pelican SVG": 1. *The Model Card Leak:* As user `rvz` pointed out, the leaked model card implies training on "user data... in accordance with Google's relevant terms of service." Given the recent terms updates, this essentially confirms that Gemini 3 is likely refined on the corpus of Gmail and Drive data that they’ve been hoarding for decades. We are cheering for a product sold back to us at a 60% markup (input costs up to $2.00/M) that was built on our own private correspondence. 2. *The "Agentic" Trap:* Everyone praising how `Gemini 3` solves complex bugs one-shot or handles entire PRs is missing the macro view. We aren’t looking at a productivity tool; we are looking at the devaluation of the software engineer. When the entry barrier drops to "can you prompt a seminal agent," the economic value of what we do on this forum evaporates. Google has successfully gamified us into feeding the very beast that will make the "14-minute human solve time" (referenced by `lairv`) irrelevant. We are optimizing for our own obsolescence while paying a monopoly rent to do it. Why is the sentiment here "Wow, cool clock widget" instead of "We just handed the keys to the kingdom to the biggest ad-tech surveillance machine in history"?
- jaapz 10mo agoGotta hand it to gemini, those are some top notch points
- BoorishBears 10mo ago
- markdog12 10mo agoI asked it to analyze my tennis serve. It was just dead wrong. For example, it said my elbow was bent. I had to show it a still image of full extension on contact, then it admitted, after reviewing again, it was wrong. Several more issues like this. It blamed it on video being difficult. Not very useful, despite the advertisements: https://x.com/sundarpichai/status/1990865172152660047 https://x.com/sundarpichai/status/1990865172152660047
- strange_quark 10mo agoI’ve never seen such a huge delta between advertised capabilities and real world experience. I’ve had a lot of very similar experiences to yours with these models where I will literally try verbatim something shown in an ad and get absolutely garbage results. Do these execs not use their own products? I don’t understand how they are even releasing this stuff.
- BoorishBears 10mo agoThe default FPS it's analyzing video at is 1, and I'm not sure the max is anywhere near enough to catch a full speed tennis serve.
- markdog12 10mo agoAh, I should have mentioned it was a slow motion video. > The default FPS it's analyzing video at is 1 Source?
- JacobAsmuth 10mo agohttps://ai.google.dev/gemini-api/docs/video-understanding#custom-frame-rate https://ai.google.dev/gemini-api/docs/video-understanding#cu... "By default 1 frame per second (FPS) is sampled from the video."
- markdog12 10mo agoOK, I just used https://gemini.google.com/app https://gemini.google.com/app, I wonder if it's the same there.
- t_minus_40 10mo agois there even a puzzle or math problem gemini 3 cant solve?
- kachapopopow 10mo agoIt's joeover for openai and antrophic. I have been using it for 3 hours now for real work and gpt-5.1 and sonnet 4.5 (thinking) does not come close. the token efficiency and context is also mindblowing... it feels like I am talking to someone who can think instead of a **rider that just agrees with everything you say and then fails doing basic changes, gpt-5.1 feels particulary slow and weak in real world applications that are larger than a few dozen files. gemini 2.5 felt really weak considering the amount of data and their proprietary TPU hardware in theory allowing them way more flexibility, but gemini 3 just works and it truly understands which is something I didn't think I'd be saying for a couple more years.
- beezlewax 10mo agoCan't wait til Gemini 4 is out!
- mparis 10mo agoI've been playing with the Gemini CLI w/ the gemini-pro-3 preview. First impressions are that its still not really ready for prime time within existing complex code bases. It does not follow instructions. The pattern I keep seeing is that I ask it to iterate on a design document. It will, but then it will immediately jump into changing source files despite explicit asks to only update the plan. It may be a gemini CLI problem more than a model problem. Also, whoever at these labs is deciding to put ASCII boxes around their inputs needs to try using their own tool for a day. People copy and paste text in terminals. Someone at Gemini clearly thought about this as they have an annoying `ctrl-s` hotkey that you need to use for some unnecessary reason.. But they then also provide the stellar experience of copying "a line of text where you then get | random pipes | in the middle of your content". Codex figured this out. Claude took a while but eventually figured it out. Google, you should also figure it out. Despite model supremacy, the products still matter.
- dr_dshiv 10mo agoMake a pelican riding a bicycle in 3d: https://gemini.google.com/share/def18e3daa39 https://gemini.google.com/share/def18e3daa39 Amazing and hilarious
- xnx 10mo agoSimilar hilarious results (one shot): https://aistudio.google.com/apps/drive/1XA4HdqQK5ixqi1jD9uMgmGUoxBeu8fFX?showPreview=true&showAssistant=true https://aistudio.google.com/apps/drive/1XA4HdqQK5ixqi1jD9uMg...
- vlmrun-admin 10mo agohttps://www.youtube.com/watch?v=cUbGVH1r_1U https://www.youtube.com/watch?v=cUbGVH1r_1U side by side comparison of gemini with other models
- vlmrun-admin 10mo agohttps://www.youtube.com/watch?v=cUbGVH1r_1U https://www.youtube.com/watch?v=cUbGVH1r_1U Everyone is talking about the release of Gemini 3. The benchmark scores are incredible. But as we know in the AI world, paper stats don't always translate to production performance on all tasks. We decided to put Gemini 3 through its paces on some standard Vision Language Model (VLM) tasks – specifically simple image detection and processing. The result? It struggled where I didn't expect it to. Surprisingly, VLM Run's Orion (https://chat.vlm.run/ https://chat.vlm.run/) significantly outperformed Gemini 3 on these specific visual tasks. While the industry chases the "biggest" model, it’s a good reminder that specialized agents like Orion are often punching way above their weight class in practical applications. Has anyone else noticed a gap between Gemini 3's benchmarks and its VLM capabilities?
- acoustics 10mo agoDon't self-promote without disclosure.
- jdthedisciple 10mo agoWhat I'd prefer over benchmarks is the answer to a simple question: What useful thing can it demonstrably do that its predecessors couldn't?
- Ridius 10mo agoKeep the bubble expanding for a few months longer.
- BugsJustFindMe 10mo agoThe Gemini AI Studio app builder (https://aistudio.google.com/apps https://aistudio.google.com/apps) refuses to generate python files. I asked it for a website, frontend and python back end, and it only gave a front end. I asked again for a python backend and it just gives repeated server errors trying to write the python files. Pretty shit experience.
- thrownaway561 10mo agoyea great.... when will I be able to have it dial a number on my google pixel? Seriously... Gemini absolutely sucks on pixel since it can't interact with the phone itself so it can't dial numbers.
- deleted 10mo ago[deleted]
- sylware 10mo agoTrained models should be able to use formal tools (for instance a logical solver, a computer?). Good. That said, I wonder if those models are still LLMs.
- falcor84 10mo agoI love it that there's a "Read AI-generated summary" button on their post about their new AI. I can only expect that the next step is something like "Have your AI read our AI's auto-generated summary", and so forth until we are all the way at Douglas Adams's Electric Monk: > The Electric Monk was a labour-saving device, like a dishwasher or a video recorder. Dishwashers washed tedious dishes for you, thus saving you the bother of washing them yourself; video recorders watched tedious television for you, thus saving you the bother of looking at it yourself. Electric Monks believed things for you, thus saving you what was becoming an increasingly onerous task, that of believing all the things the world expected you to believe. - from "Dirk Gently's Holistic Detective Agency"
- davedigerati 10mo agoExcellent reference Tried to name an AI project at work Electric Monk but too 'controversial' Had to change to Electric Mentor....
- mikepurvis 10mo agoSMBC had a pretty great take on this: https://www.smbc-comics.com/comic/summary https://www.smbc-comics.com/comic/summary
- AstroBen 10mo agoThis feels too real to laugh at
- SchemaLoad 10mo agoThere was another comic where one worker uses AI to turn their prompt in to a verbose email, then on the receiver side they use AI to turn the verbose email in to a short summary.
- drstewart 10mo agoThis one isn't a joke. 90% of documents produced at work are now AI generated, and nobody can keep up with the volume so they just summarise them with AI. What are we even doing.
- otikik 10mo ago… agentic … Meh, not interested already
- taikahessu 10mo agoBoring. Tried to explore sexuality related topics, but Alphabet is stuck in some Christianity Dark Ages. Edit: Okay, I admit I'm used to dealing with OpenAI models and it seems you have to be extra careful with wording with Gemini. Once you have right wording like "explore my own sexuality" and avoid certain words, you can get it going pretty interestingly.
- oezi 10mo agoProbably invested a couple of billion into this release (it is great as far as I can tell), but can't bring proper UI to AI Studio for long prompts and responses (e.g. it animates new text being generated even though you just return to the tab which was finished generating).
- qingcharles 10mo agoSomebody "two-shotted" Mario Bros NES in HTML: https://www.reddit.com/r/Bard/comments/1p0fene/gemini_3_the_make_mario_benchmark_has_been_beaten/ https://www.reddit.com/r/Bard/comments/1p0fene/gemini_3_the_...
- agentifysh 10mo agomy only complaint is i wish the SWE and agentic coding would have been better to justify the 1~2x premium gpt-5.1 honestly looking very comfortable given available usage limits and pricing although gpt-5.1 used from chatgpt website seems to be better for some reason Sonnet 4.5 agentic coding still holding up well and confirms my own experiences i guess my reaction to gemini 3 is a bit mixed as coding is the primary reason many of us pay $200/month for
- iib 10mo agoAs soon as I found out that this model launched, I tried giving it a problem that I have been trying to code in Lean4 (showing that quicksort preserves multiplicity). All the other frontier models I tried failed. I used the pro version and it started out well (as they all did), but it couldn't prove it. The interesting part is that it typoed the name of a tactic, spelling it "abjel" instead of "abel", even though it correctly named the concept. I didn't expect the model to make this kind of error, because they all seems so good at programming lately, and none of the other models did, although they did some other naming errors. I am sure I can get it to solve the problem with good context engineering, but it's interesting to see how they struggle with lesser represented programming languages by themselves.
- davide_benato 10mo agoI would love to see how Gemini 3 can solve this particular problem. https://lig-membres.imag.fr/benyelloul/uherbert/index.html https://lig-membres.imag.fr/benyelloul/uherbert/index.html It used to be an algorithmic game for a Microsoft student competition that ran in the mid/late 2000. The game invents a new, very simple, recursive language to move the robot (herbert) on a board, and catch all the dots while avoiding obstacles. Amazingly this clone's executable still works today on Windows machines. The interesting thing is that there is virtually no training data for this problem, and the rules of the game and the language are pretty clear and fit into a prompt. The levels can be downloaded from that website and they are text based. What I noticed last time I tried is that none of the publicly available models could solve even the most simple problem. A reasonably decent programmer would solve the easiest problems in a very short amount of time.
- keepamovin 10mo agoI don't wan't to shit on the much anticipated G3 model, but I have been using it for a complex single page task and find it underwhelming. Pro 2.5 level, beneath GPT 5.1. Maybe it's launch jitters. It struggles to produce more than 700 lines of code in a single file (aistudio). It struggles to follow instructions. Revisions omit previous gains. I feel cheated! 2.5 Pro has been clearly smarter than everything else for a long time, but now 3 seems not even as good as that, in comparison to the latest releases (5.1 etc). What is going on?
- bilsbie 10mo agoIs there a way to use this without being in the whole google ecosystem? Just make a new account or something?
- mtremsal 10mo agoIf you mean the "consumer ecosystem", then Gemini 3 should be available as an API through Google's AI Vertex platform. If you don't even want a Google Cloud account, then I think the answer is no unless they announce a partnership with an inference cloud like cerebras.
- tim333 10mo agoYou could probably do a new account. I have the odd junk google account.
- DrNosferatu 10mo agoAnyone has any idea if/when it’s coming to paid Perplexity?
- CMay 10mo agoI was sorting out the right way to handle a medical thing and Gemini 2.5 Pro was part of the way there, but it lacked some necessary information. Got the Gemini 3.0 release notification a few hours after I was looking into that, so I tried the same exact prompt and it nailed it. Great, useful, actionable information that surfaced actual issues to look out for and resolved some confusion. Helped work through the logic, norms, studies, standards, federal approvals and practices. Very good. Nice work! These things will definitely change lives.
- JohnKemeny 10mo agoThis ad was brought to you by DeepMind™. Changing people's lives.
- hekkle 10mo agoGOOGLE: "We have a new product". REALITY: It's just 3 existing products rolled into one. One of which isn't even a Google product. - Microsoft Code - Gemeni - Chrome Browser
- smarx007 10mo agoIs it coming to Google Jules?
- gabriel666smith 10mo agoInterestingly, Gemini 3 Pro Thinking clocks in at 2nd place - behind Horizon Alpha, but ahead of any actually-released model checkpoints - on the TvP Benchmark, with a score of 71.8%. This blows Sonnet 4.5 (64.8%), GPT5.1 (58.7%) and Opus 4.1 (66.3%) away. An example output: ''' The cabin was cold. It smelled of pine needles and old dust. Bryan sat on the porch with a glass of scotch balanced on his knee. It was cheap scotch. He missed the good stuff, but he wasn't in Paris anymore, and he certainly wasn't in Beverly Hills. He checked his watch. It was a bulky thing, black rubber. He checked the Glock on the table next to him. That was black, too. "I told you," Bryan said. He wasn't talking to anyone. "I told you not to go into the woods." He was thinking about Kim. He was always thinking about Kim. She was safe, probably. At a mall in Santa Monica. Or maybe a library. But it felt like she was gone. It always felt like she was gone. The trees at the edge of the clearing shifted. There wasn't any wind. The air was heavy and humid, unusual for October. A bird stopped singing. Then the crickets stopped. It got very quiet. Bryan took a sip of the scotch. It burned his throat. He liked the burn. It reminded him he was still there. Three red dots appeared on his chest. They made a small triangle. Bryan didn't look at them. He set the glass down on the railing. He picked up the phone. It was an old flip phone. He pressed a button. "I don't know who you are," Bryan said into the phone. "I don't know what you want. If you are looking for ransom, I can tell you I don't have money." The leaves rustled. A shape began to form out of the empty air. It looked like heat rising off the highway in August. It was tall. Much taller than a man. "But what I do have," Bryan said, "are a very particular set of skills." The shape solidified. It was ugly. It had skin like a lizard and dreadlocks that looked like heavy black cables. It wore a metal mask. It made a clicking sound. Click-click-click. Like a Geiger counter. "Skills I have acquired over a very long career," Bryan said. He stood up. His knees popped. He felt the ache in his lower back. He was getting too old for this. The thing roared. It wasn't a human sound. It was a wet, guttural noise. "I will look for you," Bryan said. He put the phone in his pocket. "I will find you." The creature raised a metal gauntlet. Two jagged blades slid out. They caught the moonlight. "And I will kill you," Bryan said. He raised the Glock and fired twice. Pop-pop. The creature didn't fall. It took a step forward. Green blood, thick and glowing, leaked from its shoulder. It looked like antifreeze. The creature touched the wound, then looked at its hand. It seemed surprised. Bryan didn't wait. He moved off the porch. He moved fast for a man his age. He went left, toward the woodpile. The creature fired a blast of blue light from its shoulder. It hit the porch swing. The swing exploded. Wood splinters rained down like confetti. Bryan rolled behind the logs. He breathed in. He breathed out. He could hear the heavy footsteps crunching the dry leaves. Crunch. Crunch. "You're big," Bryan whispered. "That makes you slow." He grabbed a splitting maul from the stump. It was heavy. The handle was smooth hickory. The creature came around the woodpile. It was scanning the trees. It didn't see Bryan down low. Bryan swung the maul. He put his hips into it. He swung it like he was chopping a cord of oak. The blade hit the creature in the knee. There was a snap. A loud, wet snap. The thing howled. It fell onto one knee. Bryan dropped the maul. He stepped inside the creature’s reach. He knew exactly where to hit. The throat. The armpit. The soft spot under the ribs. He hit the creature three times, hard chops with the side of his hand. It was like hitting a brick wall, but he felt something give. The creature swung its arm back. It hit Bryan in the chest. Bryan flew backward. He hit the dirt. The wind went out of him. He lay there for a second, staring up at the stars. They looked very far away. He wondered if Lenore was looking at the same stars. Probably not. She was probably sleeping. He sat up. His ribs hurt. Maybe broken. The creature was trying to stand. It was clicking again. It tapped something on its wrist. A series of red symbols started flashing. They counted down. Bryan knew a bomb when he saw one. "No," Bryan said. He tackled the thing. He didn't think about it. He just did it. He grabbed the creature’s arm. He twisted the wrist mechanism. He’d seen something like it in Baghdad once. Or maybe Istanbul. The memories ran together now. He ripped the gauntlet loose. Wires sparked. He threw it as hard as he could into the darkness of the woods. Three seconds later, there was a flash. A boom. A shockwave that shook the pine needles from the trees. Silence came back. The creature lay on the ground. It was breathing shallowly. The green blood was pooling under it. It took off its mask. The face was hideous. Mandibles. Beady eyes. It looked at Bryan. It said something, a garbled copy of Bryan's own voice. "...good luck..." Then it died. It just stopped. Bryan stood up. He dusted off his pants. He walked back to the porch. The swing was gone. The railing was scorched. His glass of scotch was still sitting there, untouched. The ice hadn't even melted. He picked it up. He took a drink. It still tasted cheap. He took his phone out and looked at it. No service. "Well," he said. He went inside the cabin and locked the door. He sat on the couch and waited for the sun to come up. He hoped Kim would call. He really hoped she would call. '''
- deleted 10mo ago[deleted]
- visioninmyblood 10mo agoReally exciting results on paper. But truly interesting to see what data this has been trained on. There is a thin line between accuracy improvements and the data used from users. Hope the data used to train was obtained with consent from the creators
- raffkede 10mo agoSeems to be the first model that one-shots my secret benchmark about nested SQLite and it did it in 30s,
- osn9363739 10mo agoOut of interest. Does it one shot it every time?
- raffkede 10mo agoWill try again just tried once in the phone a few hours ago, other models were able to do quite a lot but usually missing some stuff this time it managed nested navigation quite well, lot of stuff missing for sure I just tested the basics with the play button in AI studio
- osn9363739 10mo agoIt seems to be that first impression that makes all the difference. Especially with the randomness that comes with llms in general. which maybe explains the 'wow this is so much better' vs the 'this is no better than xxx' commments littered throughout this whole parent post.
- hamasho 10mo agoI just googled latest LLM models and this page appears at the top. It looks like Gemini Pro 3 can score 102% in high school math tests.
- jpkw 10mo agoHoping someone here may know the answer to this, but do any of the benchmarks that exist currently account for false answers in any meaningful way, other than it would in a typical test (ie, if I give any answer at all it is better than saying "I don't know" as the answer I give at least has a chance of being correct(which in the real world is bad))? I want an LLM that tells me when it doesn't know something. If it gives me an accurate response 90% of the time and an inaccurate one 10% of the time, it is less useful than one that gives me an accurate answer 10% of the time and tells me "I don't know" the other 90%.
- terandle 10mo agohttps://artificialanalysis.ai/evaluations/omniscience https://artificialanalysis.ai/evaluations/omniscience
- rocqua 10mo agoThose numbers are too good to expect. If 90% right 10% wrong is the baseline would you take as an improvement: - 80% right 18% I don't know 2% wrong - 50%/48%/2% - 10%/90%/0% - 80%/15%/5% The general point being that to reduce wrong answers you will need to accept some reduction in right answers if you want the change to only be made through trade-offs. Otherwise you just say "I'd like a better system" and that is rather obvious. Personally I'd take like 70/27/3. Presuming the 70% of right answers aren't all the trivial questions.
- energy123 10mo agoOpenAI uses SimpleQA to assess hallucinations
- davidpolberger 10mo agoThis is wild. I gave it some legacy XML describing a formula-driven calculator app, and it produced a working web app in under a minute: https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%5B%2218nsOyEA4Y-oypiASRQ5NjTZpAgNAa2oE%22%5D,%22action%22:%22open%22,%22userId%22:%22110718778558981006204%22,%22resourceKeys%22:%7B%7D%7D&usp=sharing https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%... I spent years building a compiler that takes our custom XML format and generates an app for Android or Java Swing. Gemini pulled off the same feat in under a minute, with no explanation of the format. The XML is fairly self-explanatory, but still. I tried doing the same with Lovable, but the resulting app wouldn't work properly, and I burned through my credits fast while trying to nudge it into a usable state. This was on another level.
- zarzavat 10mo agoThis is exactly the kind of task that LLMs are good at. They are good at transforming one format to another. They are good at boilerplate. They are bad at deciding requirements by themselves. They are bad at original research, for example developing a new algorithm.
- codespin 10mo ago> They are good at transforming one format to another. They are good at boilerplate. You just described 90% of coding
- taco_emoji 10mo agoMaybe 90% of the actual typing part of coding, but not 90% of the JOB of coding.
- nwienert 10mo agoThey’re bad at 90% of coding, but for other reasons. That said if you babysit them incessantly they can help you move a bit faster through some of it.
- oblio 10mo ago
- cognitive-gl 10mo agoWild
- ogig 10mo agoI just gave it a short description of a small game I had an idea for. It was 7 sentences. It pretty much nailed a working prototype, using React, clean css, Typescript and state management. It event implemented a Gemini query using the API for strategic analysis given a game state. I'm more than impressed, I'm terrified. Seriously thinking of a career change.
- brcmthrowaway 10mo agoTo what?
- apparent 10mo agoVC (vibe coding).
- osn9363739 10mo agoCan you share the code?
- rkozik1989 10mo agoNo because this story didn't happen.
- ogig 10mo agohttps://ai.studio/apps/drive/1E-aYovHHoY8jrF6bsl_AZ8VszIN66Nap https://ai.studio/apps/drive/1E-aYovHHoY8jrF6bsl_AZ8VszIN66N... The initial prompt was, in case people doesn't want to log in: Make a turn based chess like game. Instead of normal chess board use an hexagonal grid. Make the board diagonal shaped. Instead of traditional chess pieces we are going to use spaceship designs. Each spaceship has unique abilities that influence the board or their own skill. For 2 players, turn based. Show me what you got.
- WhyOhWhyQ 10mo agoI just spent 12 hours a day vibe coding for a month and a half with Claude (which has equal swe benchmarks at gemini 3). I started out terrified but eventually I realized that these are just remarkably far away from actually replacing a real software engineer. For prototypes they're amazing, but when you're just straight vibe coding you get stuck in a hell where you don't want to or can't efficiently really check what's going on under the hood but it's not really doing the thing you want. Basically these tools can you you to a 100k LOC project without much effort, but it's not going to be a serious product. A serious product requires understanding still.
- King-Aaron 10mo ago> it’s been incredible to see how much people love it. AI Overviews now have 2 billion users every month "Incredible"! When they insert it into literally every google request without an option to disable it. How incredibly shocking so many people use it.
- zen_boy 10mo agoIs the "thinking" dropdown option on gemini.google.com what the blog post refers to as Deep Think?
- maczwei 10mo agoentity.ts is in types/entity.ts .it cant grasp that it should import it like "../types/entity" and instead it always writes "../types" i am using the https://aistudio.google.com/apps https://aistudio.google.com/apps
- kanodiaayush 10mo agoI don't really understand the amount of ongoing negativity in the comments. This is not the first time a product has been near copied, and the experience for me is far superior to code in a terminal. It comes with improvements even though imperfect, and I'm excited for those! I've long wanted the ability to comment on code diffs instead of just writing things back down in chat. And I'm excited for the quality of gemini 3.0 pro; although I'm running into rate limits. I can already tell its something I'm going to try out a lot!
- rvnx 10mo agoIt's not really good for real-life programming though, it invents lot of imaginary things, cannot respect its own instructions, forgets basic things (variable is called "bananaDance", then claims it is "bananadance", then later on "bananaDance" again). It is good at writing something from scratch (like spitting out its training set). Claude is still superior for programming and debugging. Gemini is better at daily life questions and creative writing.
- kanodiaayush 10mo agoyeah testing it out! good to know the above. My feel also is that claude is better so far.
- rvnx 10mo agoIt's not bad at all though, but it needs lot a baby-sitting like "try again, try this, try that, are you sure that it is correct ?" For example, in a basic python script that uses os.path.exists, it forgets the basic "import os", and then, "I apologize for the oversight".
- kanodiaayush 10mo agoSimilar stuff my end; I'm coding up a complex feature - Claude would have taken fewer interventions on my part, and would have been non buggy right off the bat. But apart from that the experience is comparable.
- chiragsrvstv 10mo agoWaiting for google to nuke this as well just like 2.5pro
- eterm 10mo ago> It seems there's a date conflict. The prompt claims it's 2025, but my internal clock says otherwise. > I'm now zeroing in on the temporal aspect. Examining the search snippets reveals dates like "2025-10-27," suggesting a future context relative to 2024. My initial suspicion was that the system time was simply misaligned, but the consistent appearance of future dates strengthens the argument that the prompt's implied "present" is indeed 2025. I am now treating the provided timestamps as accurate for a simulated 2025. It is probable, however, that the user meant 2024. Um, huh? It's found search results for October 2025, but this has led it to believe it's in a simulated future, not a real one?
- pclark 10mo agoI just want Gemini to access ALL my Google Calendars, not just the primary one. If they supported this I would be all in on Gemini. Does no one else want this?
- petesergeant 10mo agoStill insists the G7 photo[0] is doctored, and comes up with wilder and wilder "evidence" to support that claim, before getting increasingly aggressive. 0: https://en.wikipedia.org/wiki/51st_G7_summit#/media/File:Prime_Minister_Keir_Starmer_attends_the_G7_Summit_in_Canada_(54594644396).jpg https://en.wikipedia.org/wiki/51st_G7_summit#/media/File:Pri...
- kmeisthax 10mo agoThe most devastating news out of this announcement is that Vending-Bench 2 came out and it has significantly less clanker[0] meltdowns than the first one. I mean, seriously? Not even one run where the model tried to stock goods that hadn't arrived yet, only for it to eventually try and fail to shut down the business, and then e-mail the FBI about the $2 daily fee being deducted from the bot? [0] Fake racial slur for a robot, LLM chatbot, or other automated system
- auggierose 10mo ago> Gemini 3 is the best vibe coding and agentic coding model we’ve ever built Google goes full Apple...
- elcapithanos 10mo ago> AI overviews now have 2 billion users every month More like 2 billion hostages
- lofaszvanitt 10mo agoA tad bit better, still has the same issues regarding unpacking and understanding complex prompts. I have a test of mine and now it performs a bit better, but still, it has zero understanding what is happening and for why. Gemini is the best of the best model out there, but with complex problems it just goes down the drain :(.
- lofaszvanitt 10mo agoOh that corpulent fella with glasses who talks in the video. Look how good mannered he is, he can't hurt anyone. But Google still takes away all your data and you will be forced out of your job.
- energy123 10mo agoWith the $20/m subscription, do we get it on "Low" or "High" thinking level?
- primaprashant 10mo agoCreated a summary of comments from this thread about 15 hours after it had been posted and had 814 comments with gemini-3-pro and gpt-5.1 using this script [1]: - gemini-3-pro summary: https://gist.github.com/primaprashant/948c5b0f89f1d5bc919f903da5e5d94f https://gist.github.com/primaprashant/948c5b0f89f1d5bc919f90... - gpt-5.1 summary: https://gist.github.com/primaprashant/3786f3833043d8dcccae4bfd4ff9f4a7 https://gist.github.com/primaprashant/3786f3833043d8dcccae4b... Summary from GPT 5.1 is significantly longer and more verbose compared to Gemini 3 Pro (13,129 output tokens vs 3,776). Gemini 3 summary seems more readable, however, GPT 5.1 one has interesting insights missed by Gemini. Last time I did this comparison at the time of GPT 5 release [2], the summary from Gemini 2.5 Pro was way better and readable than the GPT 5 one. This time the readability of Gemini 3 summary still seems great while GPT 5.1 feels a bit more improved but not there quite yet. [1]: https://gist.github.com/primaprashant/f181ed685ae563fd06c49d3d49a8dd9b https://gist.github.com/primaprashant/f181ed685ae563fd06c49d... [2]: https://news.ycombinator.com/item?id=44835029 https://news.ycombinator.com/item?id=44835029
- misja111 10mo agoI asked Gemini to solve today's Countle puzzle (https://www.countle.org/ https://www.countle.org/). It got stuck while iterating randomly trying to find a solution. While I'm writing this it has been trying already for 5 minutes and the web page has become unresponsive. I also asked it for the best play when in backgammon opponent rolls 6-1 (plays 13/7 8/7) and you roll 5-1. It starts alright with mentioning a good move (13/8 6/5) but continues to hallucinate with several alternative but illegal moves. I'm not too impressed.
- iamA_Austin 10mo agoit started with OpenAI and Google took the competition damn seriously.
- pk-protect-ai 10mo agoIt is pointless to ask an LLM to draw an ASCII unicorn these days. Gemini 3 draws one of these (depending on the prompt): https://www.ascii-art.de/ascii/uvw/unicorn.txt https://www.ascii-art.de/ascii/uvw/unicorn.txt However, it is amazing how far spatial comprehension has improved in multimodal models. I'm not sure the below would be properly displayed on HN; you'll probably need to cut and paste it into a text editor. Prompt: Draw me an ASCII world map with tags or markings for the areas and special places. Temperature: 1.85 Top-P 0.98 Answer: Edit (replaced with URL) https://justpaste.it/kpow3 https://justpaste.it/kpow3
- nprateem 10mo agoOMG they've obviously had a major breakthrough because now it can reply to questions with actual answers instead of shit blog posts.
- rubymamis 10mo agoI gave it the task to recreate StackView.qml to be feel more native on iOS and it failed - like all other models... Prompt: Instead of the current StackView, I want you to implement a new StackView that will have a similar api with the differences that: 1. It automatically handles swiping to the previous page/item. If not mirrored, it should detect swiping from the left edge, if mirrored it should detect from the right edge. It's important that swiping will be responsive - that is, that the previous item will be seen under the current item when swiping - the same way it's being handled on iOS applications. You should also add to the api the option for the swipe to be detected not just from the edge, but from anywhere on the item, with the same behavior. If swiping is released from x% of current item not in view anymore than we should animate and move to the previous item. If it's a small percentage we should animate the current page to get back to its place as nothing happened. 2. The current page transitions are horrible and look nothing like native iOS transitions. Please make the transitions feel the same.
- bluecalm 10mo agoI've asked it (thinking 3) about the difference between Plus and Pro plans. First it thought I am asking for comparison between Gemini and ChatGPT as it claimed there is no "Plus" plan on Gemini. After I insisted I am on this very plan right now it apologized and told me it in fact exists. Then it told me the difference is that I got access to newer models with the Pro subscription. That is despite Google's own plan comparison page showing I get access to the Gemini 3 on both plans. It also told me that on Plus I am most likely using "Flash" model. There is no "Flash" model in the dropdown to choose from. There is only "Fast" and "Thinking". It then told me "Fast" is just renamed Flash and it likely uses Gemini 2.5. On the product comparison page there is nothing about 2.5, it only mentions version 3 for both Plus and Pro plans. Of course on the dropdown menu it's impossible to see which model it is really using. How can a normal person understand their products when their own super advanced thinking/reasoning model that took months to train on world's most advanced hardware can't? It's amazing to me they don't see it as an epic failure in communication and marketing.
- jacky2wong 10mo agoWhat I loved about this release was that it was hyped up by a polymarket leak with insider trading - NOT with nonsensical feel the AGI hype. Great model that's pushed the frontier of spatial reasoning by a long shot.
- taf2 10mo agoI had asked earlier in the day for gpt 5.1 high to refactor my apex visualforce page into a lightning component and it really didn’t do much here - Gemini 3 pro crushed this task… very promising
- gigatexal 10mo agoHow does it do in coding tasks? I’ve been absolutely spoiled by Claude sonnet 4.5 thinking.
- abixb 10mo agoOkay, Gemini 3.0 Pro has officially surpassed Claude 4.5 (and GPT-5.1) as the top ranked model based on my private evals (multimodal reasoning w/ images/audio files and solving complex Caesar/transposition ciphers, etc.). Claude 4.5 solved it as well (the Caesar/transposition ciphers), but Gemini 3.0 Pro's method and approach was a lot more elegant. Just my $0.02.
- tim333 10mo agoHassabis interview on Gemini 3, with Hard Fork (nyt podcast), also Josh Woodward https://youtu.be/rq-2i1blAlU?t=428 https://youtu.be/rq-2i1blAlU?t=428 Some points - Good at vibe coding 10:30 - step change where it's actually useful AGI still 5-10 years. Needs reasoning, memory, world models. Is it a bubble? - Partly 22:00 What's fun to do with Gemini to show the relatives? Suggested taking a selfie with the app and having it edit. 24:00 (I tried and said make me younger. Worked pretty well.) Also interesting - apparently they are doing an agent to go through your email inbox and propose replies automatically 4:00. I could see that getting some use.
- erikpukinskis 10mo ago> Needs reasoning, memory, world models. Is that all? So they just need to invent: 1. Thought 2. A mechanism for efficiently encoding and decoding arbitrary percepts 3. A formal model of the world And then the existing large language models can handle the rest. Yep, 5 years and a hundred billion dollars or so should do the trick.
- mark_l_watson 10mo agoI had a fantastic ‘first result’ with Gemini 3 but a few people on social media I respect didn’t. Key takeaway is to do your own testing with your use cases. I feel like I am now officially biased re: LLM infrastructure: I am retired, doing personal research and writing, and I decided months ago to drop OpenAI and Anthropic infrastructure and just use Google to get stuff done - except I still budget about two hours a week to experiment with local models and Chinese models’ APIs.
- taf2 10mo agoI just wish gemini could write well formatted code. I do like the solutions it comes up to and I know I can use a linter/formatter tool - but it would just be nice if when I openned gemini (cli) up and asked it to write a feature it didn't mix up the indenting so badly... somehow codex and claude both get this without any trouble...
- AbstractH24 10mo agoCan someone ELI5 what the difference between AI Studio, Antigravity, and Colab is?
- simlevesque 10mo agoAi studio is a web chat. Antigravity is an IDE you install. Colab is a place to run notebooks in the cloud.
- AbstractH24 10mo agoColab has significant Gemini functionality built in. How isn't it a combination of the first two? Thanks for sorting all this out! Still exploring the first two, so I really don't know.
- WhyOhWhyQ 10mo agoWhy doesn't this spell the death of OpenAI? Maybe someone with a better business sense can explain, but here's what I'm seeing: OpenAI is going for the consumer-grade AI market, as opposed to a company like Anthropic making a specialized developer tool. Google can inject their AI tool in front of everybody in the world, and already have with Google AI search. All of these models are just going to reach parity eventually, but Google is burning cash compared to OpenAI burning debt. It seems like for consumer-grade purposes, AI use will just be free sooner or later (DeepSeek is free, Google AI search is free, students can get Gemini Pro for free for a year already). So all I'm seeing that OpenAI has is Sora, which seems like a business loser though I don't really understand it, and also ChatGPT seems to own the market of people roleplaying with chat bots as companions (which doesn't really seem like a multi-trillion dollar business but I could be wrong).
- nextworddev 10mo agoYep. Except OpenAI is mainly burning LP money (saudis, softbank, pension funds)
- thingsilearned 10mo agoI love that the recipe example is still being used as one of the main promising use cases for computers and now AGI. One day hopefully computers will solve that pressing problem...
- espeed 10mo agoI paid for Gemini Pro. Am I getting Gemini 3 Pro (https://gemini.google.com https://gemini.google.com)? "To be precise: You are currently interacting with Gemini 1.5 Pro." https://x.com/espeed/status/1991333475098718601 https://x.com/espeed/status/1991333475098718601
- decide1000 10mo agoWe hire a developer to build parsers for a complicated file format. It takes a week per parser. Gemini 3 is the first LLM that is able to create a parser from scratch, and it does it very well. Within a minute, 1-shot-right. I am blown away.
- Frannky 10mo agoI tried it on a landing page. Very, very impressive.
- VladimiOrlovsky 10mo agoimport decimal def solve_kangaroo_limit(): # Set precision to handle the "digits different from six" requirement decimal.getcontext().prec = 50 # For U(0,1), H(x) approaches 2x + 2/3 very rapidly (exponential decay of error) # At x = 10^6, the value is indistinguishable from the asymptote x = 10**6 limit_value = decimal.Decimal(2) * x + decimal.Decimal(2) / decimal.Decimal(3) print(f"H({x}) ≈ {limit_value}") # Output: 2000000.66666666666666666666... if __name__ == "__main__": solve_kangaroo_limit() ....p.s. for airheads=idiots: """decimal.Decimal(2) / decimal.Decimal(3)""" == 0.6666666666666666666666666666666666666666666666666666666666666666666666666 ... This is your Fukingly 'smart' computer???
- VladimiOrlovsky 10mo agop.s. This is not on singular error, this is a good example of fundamental problems ... and a lot of lies. End of conversation.