7 ms·
At this point in time I start to believe OAI is very much behind on the models race and it can't be reversed Image model they have released is much worse than
by Obertr 9mo ago
At this point in time I start to believe OAI is very much behind on the models race and it can't be reversed
Image model they have released is much worse than nano banana pro, ghibli moment did not happen
Their GPT 5.2 is obviously overfit on benchmarks as a consensus of many developers and friends I know. So Opus 4.5 is staying on top when it comes to coding
The weight of the ads money from google and general direction + founder sense of Brin brought the google massive giant back to life.
None of my companies workflow run on OAI GPT right now. Even though we love their agent SDK, after claude agent SDK it feels like peanuts.
- random9749832 9mo agoThis is obviously trained on Pro 3 outputs for benchmaxxing.
- CuriouslyC 9mo agoNot trained on pro, distilled from it.
- viraptor 9mo agoWhat do you think distilled means...?
- CuriouslyC 9mo agoIt's good to keep the language clear, because you could pretrain/sft on outputs (as many labs do), which is not the same thing.
- NitpickLawyer 9mo ago> for benchmaxxing. Out of all the big4 labs, google is the last I'd suspect of benchmaxxing. Their models have generally underbenched and overdelivered in real world tasks, for me, ever since 2.5 pro came out.
- dieortin 9mo agoIs there anything pointing to Brin having anything to do with Google’s turnaround in AI? I hear a lot of people saying this, but no one explaining why they do
- ryoshu 9mo agoIf he's having an impact it's because he can break through the bureaucracy. He's not trying to protect a fiefdom.
- novok 9mo agoIn organizations, everyone's existence and position is politically supported by their internal peers around their level. Even google's & microsoft's current CEOs are supported by their group of co-executives and other key players. The fact that both have agreeable personalities is not a mistake! They both need to keep that balance to stay in power, and that means not destroying or disrupting your peer's current positions. Everything is effectively decided by informal committee. Founders are special, because they are not beholden to this social support network to stay in power and founders have a mythos that socially supports their actions beyond their pure power position. The only others they are beholden too are their co-founders, and in some cases major investor groups. This gives them the ability to disregard this social balance because they are not dependent on it to stay on power. Their power source is external to the organization, while everyone else is internal to it. This gives them a very special "do something" ability that nobody else has. It can lead to failures (zuck & occulus, snapchat spectacles) or successes (steve jobs, gemini AI), but either way, it allows them to actually "do something".
- JumpCrisscross 9mo ago> Founders are special, because they are not beholden to this social support network to stay in power Of course they are. Founders get fired all the time. As often as non-founder CEOs purge competition from their peers. > The only others they are beholden too are their co-founders, and in some cases major investor groups This describes very few successful executives. You can have your co-founders and investors on board, if your talent and customers hate you, they’ll fuck off.
- avazhi 9mo ago"At this point in time I start to believe OAI is very much behind on the models race and it can't be reversed" This has been true for at least 4 months and yeah, based on how these things scale and also Google's capital + in-house hardware advantages, it's probably insurmountable.
- mmaunder 9mo agoYeah the only thing standing in Google's way is Google. And it's the easy stuff, like sensible billing models, easy to use docs and consoles that make sense and don't require 20 hours to learn/navigate, and then just the slew of bugs in Gemini CLI that are basic usability and model API interaction things. The only differentiator that OpenAI still has is polish. Edit: And just to add an example: openAI's Codex CLI billing is easy for me. I just sign up for the base package, and then add extra credits which I automatically use once I'm through my weekly allowance. With Gemini CLI I'm using my oauth account, and then having to rotate API keys once I've used that up. Also, Gemini CLI loves spewing out its own chain of thought when it gets into a weird state. Also Gemini CLI has an insane bias to action that is almost insurmountable. DO NOT START THE NEXT STAGE still has it starting the next stage. Also Gemini CLI has been terrible at visibility on what it's actually doing at each step - although that seems a bit improved with this new model today.
- mips_avatar 9mo agoI'd be curious how many people use openrouter byok just to avoid figuring out the cloud consoles for gcp/azure.
- baq 9mo agoGPT 5.2 is actually getting me better outputs than Opus 4.5 on very complex reviews (on high, I never use less) - but the speed makes Opus the default for 95% of use cases.
- raincole 9mo agoThat's a quite sensationalized view. Ghibli moment was only about half a year ago. At that moment, OpenAI was so far ahead in terms of image editing. Now it's behind for a few months and "it can't be reversed"?
- Obertr 9mo agoCheck the size and budget of Google iniatives. It’s unlimited
- akie 9mo agoGoogle basically has unlimited budget and unlimited data. If they're ahead now, which I believe they are, they'll be very very difficult to catch.
- BoredPositron 9mo agoThe Ghibli moment was an influencer fad not real advancement.
- encroach 9mo agoOAI's latest image model outperforms Google's in LMArena in both image generation and image editing. So even though some people may prefer nano banana pro in their own anecdotal tests, the average person prefers GPT image 1.5 in blind evaluations. https://lmarena.ai/leaderboard/text-to-image https://lmarena.ai/leaderboard/text-to-image https://lmarena.ai/leaderboard/image-edit https://lmarena.ai/leaderboard/image-edit
- Obertr 9mo agoAdd This to Gemini distribution which is being adcertised by Google in all of their products, and average Joe will pick the sneakers at the shelf near the checkout rather than healthier option in the back
- encroach 9mo agoThat's not how the arena works. The evaluation is blind so Google's advertising/integration has no effect on the results.
- gdhkgdhkvff 9mo agoThose darn sneakers are just too delicious!
- raincole 9mo ago...and what does this have to do with the comment you replied to? Did you reply to the wrong person or you were just stating unrelated factoids?
- int32_64 9mo agoIs there a "good enough" endgame for LLMs and AI where benchmarks stop mattering because end users don't notice or care? In such a scenario brand would matter more than the best tech, and OpenAI is way out in front in brand recognition.
- holler 9mo agothis. I don't know any non-tech people who use anything other than chatgpt. On a similar note, I've wondered why Amazon doesn't make a chatgpt-like app with their latest Alexa+ makeover, seems like a missed opportunity. The Alexa app has a feature to talk to the LLM in chat mode, but the overall app is geared towards managing devices.
- Obertr 9mo agoMost of Europe if full of Gemini ads, my parents use Gemini because it is free and it popped up in YouTube ad before the video Just go outside the bubble plus take a bit older people
- ewoodrich 9mo agoYeah my parents never really cared enough to explore ChatGPT despite hearing about it 10 times a day in news/media for the last few years. But recently my mom started using Google's AI Search mode after first trying it while doing research for house hunting and my dad uses the Gemini app for occasional questions/identifying parts and stuff (he has always loved Google Lens so those sort of interactive multimedia features are the main pull vs plain text chatbot conversations). They are both Android/Google Search users so all it really took was "sure I guess I'll try that" in response to a nudge from Google. For me personally I have subscriptions to Claude/ChatGPT/Gemini for coding but use Gemini for 90% of chatbot questions. Eventually I'll cancel some of them but will probably keep Gemini regardless because I like having the extra storage with my Google One plan bundle. Google having a pre-existing platform/ecosystem is a huge advantage imo.
- macNchz 9mo agoGoogle has great distribution to be able to just put Gemini in front of people who are already using their many other popular services. ChatGPT definitely came out of the gate with a big lead on name recognition, but I have been surprised to hear various non-techy friends talking about using Gemini recently, I think for many of them just because they have access at work through their Workspace accounts.
- yieldcrv 9mo agothe trend I've seen is that none of these companies are behind in concept and theory, they are just spending longer intervals baking a more superior foundational model so they get lapped a few times and then drop a fantastic new model out of nowhere the same is going to happen to Google again, Anthropic again, OpenAI again, Meta again, etc they're all shuffling the same talent around, its California, that's how it goes, the companies have the same institutional knowledge - at least regarding their consumer facing options
- GenerWork 9mo agoI'm actually liking 5.2 in Codex. It's able to take my instructions, do a good job at planning out the implementation, and will ask me relevant questions around interactions and functionality. It also gives me more tokens than Claude for the same price. Now, I'm trying to white label something that I made in Figma so my use case is a lot different from the average person on this site, but so far it's my go to and I don't see any reason at this time to switch.
- gpt5 9mo agoI've noticed when it comes to evaluating AI models, most people simply don't ask difficult enough questions. So everything is good enough, and the preference comes down to speed and style. It's when it becomes difficult, like in the coding case that you mentioned, that we can see the OpenAI still has the lead. The same is true for the image model, prompt adherence is significantly better than Nano Banana. Especially at more complex queries.
- fellowniusmonk 9mo agoI have a very complex set of logic puzzles I run through my own tests. My logic test and trying to get an agent to develop a certain type of ** implementation (that is published and thus the model is trained on to some limited extent) really stress test models, 5.2 is a complete failure of overfitting. Really really bad in an unrecoverable infinite loop way. It helps when you have existing working code that you know a model can't be trained on. It doesn't actually evaluate the working code it just assumes it's wrong and starts trying to re-write it as a different type of **. Even linking it to the explanation and the git repo of the reference implementation it still persists in trying to force a different **. This is the worst model since pre o3. Just terrible.
- GenerWork 9mo agoI'd argue that 5.2 just barely squeaks past Sonnet 4.5 at this point. Before this was released, 4.5 absolutely beat Codex 5.1 Medium and could pretty much oneshot UI items as long as I didn't try to create too many new things at once.
- 9mo ago
- louiereederson 9mo agoi think the most important part of google vs openai is slowing usage of consumer LLMs. people focus on gemini's growth, but overall LLM MAUs and time spent is stabilizing. in aggregate it looks like a complete s-curve. you can kind of see it in the table in the link below but more obvious when you have the sensortower data for both MAUs and time spent. the reason this matters is slowing velocity raises the risk of featurization, which undermines LLMs as a category in consumer. cost efficiency of the flash models reinforces this as google can embed LLM functionality into search (noting search-like is probably 50% of chatgpt usage per their july user study). i think model capability was saturated for the average consumer use case months ago, if not longer, so distribution is really what matters, and search dwarfs LLMs in this respect. https://techcrunch.com/2025/12/05/chatgpts-user-growth-has-slowed-report-finds/ https://techcrunch.com/2025/12/05/chatgpts-user-growth-has-s...
- JumpCrisscross 9mo ago> I start to believe OAI is very much behind Kara Swisher recently compared OpenAI to Netscape.
- Andrex 9mo agoOuch. Maybe we'll get some awesome FOSS tech out of its ashes?
- JumpCrisscross 9mo agoWe’ll get a bail-out and then a massive data-centre and energy-production build-out.
- nightski 9mo agoGoogle has incredible tech. The problem is and always has been their products. Not only are they generally designed to be anti-consumer, but they go out of their way to make it as hard as possible. The debacle with Antigravity exfiltrating data is just one of countless.
- novok 9mo agoThe Antigravity case feels like a pure bug and them rushing to market. They had a bunch of other bugs showing that. That is not anti-consumer or making it difficult.
- deleted 9mo ago[deleted]
- aswegs8 9mo agoNot sure why they just not replicate the workflow that nano banana pro uses. It lets the thinking model generate a detailed description and then renders that image. When I use ChatGPT thinking model and render an image I also get pretty good results. It's not as creative or flexible as nano banana pro, but it produces really useful results.