6 ms·
They didnt include chatgpt in the comparison chart. That tells a lot
by algoth1 3mo ago
They didnt include chatgpt in the comparison chart. That tells a lot
- minimaxir 3mo agoThat is fair to point out. For those who don't know, ChatGPT Image 2 has an absurd ELO of 1387; compared to the #2 model at 1273, it's over 100 points higher (https://arena.ai/leaderboard/text-to-image https://arena.ai/leaderboard/text-to-image). The tradeoff is latency, and ChatGPT Image 2 at High is...slow (~2 minutes at 1024x1024). In both cases it would have skewed the charts here to uselessness. I want to do a writeup on ChatGPT Image 2 but at this point I don't think people care about nuanced image generation anymore...even though ChatGPT Image 2 crushes all my existing tests.
- shmolyneaux 3mo agoI definitely appreciated your post about Nano Banana Pro. It's also a genuinely useful time-capsule for how these systems evolve and where they fall short. I've preferred the output of ChatGPT Image 2. I think a post would be very helpful for folks to see what they're missing.
- vunderba 3mo agoThat arena leaderboard has some questionable results. Anyone who's used these models would know that ranking HiDream above Krea2 is a pretty hot take. Many of these ELO comparative tests (ArtificialAnalysis is guilty as hell on this as well) also have other problems such as a considerable number of "amateur judges" tending to prioritize aesthetics over actual instruction-following given the prompt. Also (less a critique of Arena.AI necessarily), but the MAI models are so incredibly locked down (e.g. censored) as to be functionally useless. I have a sneaking suspicion its fallout from Tay. https://en.wikipedia.org/wiki/Tay_(chatbot) https://en.wikipedia.org/wiki/Tay_(chatbot)
- revolvingthrow 3mo agoWhile I have no experience with it personally (no interest in image gen) my aunt was raving about current chatgpt image model for "restoring" / working with old photos - sharpening, changing some small details like ill-fitting background. It takes her a bunch of prompting but eventually she gets things just right. In comparison, current gemini output (supposedly) tends to be subtly off, details aren’t quite right, proportions are subtly changed etc. This is purely about generating images with people in them, I don’t think she’s doing any logic puzzles with gotchas and specific alignments of differently colored blocks and whatnot
- HDBaseT 3mo agoChatGPT Image 2 is genuinely insane. I was surprised the internet didn't have a meltdown when it released.
- sokz 3mo agoChatGPT image is a lot more aesthetically coherent when I do a layman's test between Gemini and ChatGPT (both through their chat interface). It just breaks down in subsequent editing (around 3-4 editing prompts). That is to say Gemini also does the same, but I felt it degrades more gracefully in subsequent editing runs.
- pablonaj 3mo agoPlease do write about it, there are still definitely people interested. I have also have noticed that GPT Image 2 is very good and has a great cost/result ratio compared to other models, specially using the low version which can cost 0.01 and is usually good enough for many use cases. I completely replaced Nano Banana for most of my workflows because I'm using API and the cost adds up. I still haven't tried Nano Banana 2 Lite but the price may be hard to justify, although the speed bump sounds good.
- johny115 3mo agoand also left out Nano Banana Pro, which in my opinion is still better in many cases then NB2 ... even on the site arena.ai, it's still best at editing
- vunderba 3mo agoYeah given Google's order of progression NB NB Pro NB 2 A lot of us are still expecting Google to drop a Pro version of NB 2 but that hasn't happened yet...