5 ms·
Looks like this model is meaningfully less good than gpt-image-2. Arena.ai score is 1263 vs 1380. https://arena.ai/leaderboard/text-to-image https://arena.ai/l
by sinak 1mo ago
Looks like this model is meaningfully less good than gpt-image-2. Arena.ai score is 1263 vs 1380.
https://arena.ai/leaderboard/text-to-image https://arena.ai/leaderboard/text-to-image
- minimaxir 1mo agoEverything is less good than gpt-2-image and I suspect that will be the case for awhile, until potentially Nano Banana Pro 2. However, cost is significantly lower in this case. A Pro image here is $0.04, a gpt-2-image high is $0.21 and lower resolution.
- Invictus0 1mo agoGpt 2 image is unlimited on a chatgpt pro plan
- dragonwriter 1mo agoWhich makes it a clear price win if you already subscribe to Pro for some other reason, otherwise the price crossover where ChatGPT Pro unlimited beats the per-image pricing cited here is 80+ images/day.
- svachalek 1mo agogpt-2-image has such a yellow tinge though. It stands out badly on nearly any screen. The image quality is great otherwise, but I generally prefer nano banana for its color balance.
- vunderba 1mo agoTrue but it’s at least addressable with some basic tone-mapping changes. It’s easier to correct an issue like this than to deal with an image that simply doesn’t follow your prompt. Here's a quick trend of the "piss filter" in the gpt image series: https://imgpb.com/vCZidh https://imgpb.com/vCZidh
- CamperBob2 1mo agoNot a great test case, considering how much of the artwork in-distribution for that type of image will have age-yellowed lacquer.
- vunderba 1mo agoOh yeah, that’s a good point I’ll generate some images of a theoretical earth with a solar system bathed in the light of a K-type orange dwarf star instead. :) I have some other examples of it as well that aren't going to be potentially contaminated by that late 18th century / early 19th century tintype-esque training data. https://imgpb.com/vTUHo https://imgpb.com/vTUHo
- deleted 1mo ago[deleted]
- DANmode 1mo agoLess than 10% is meaningful? If these were processors, I wouldn’t spend another hundred on the faster one…
- minimaxir 1mo agoThe scores are pairwise ELO rankings, not simple metric scores.
- vunderba 1mo agoAs I’ve said before, I wouldn’t put a lot of stock in Arena’s scoring system. They have MAI Image 2.5 ranked above Gemini Nano Banana Pro, and maybe that’s true on paper but good luck using it. Microsoft’s censorship makes Google feel like the wild west by comparison. They also have Meta’s Muse Image ranked above NB Pro, which is just patently absurd. In my own GenAI benchmark it only managed a lackluster 7 passes out of 15. Even the open‑weight Ideogram 4 scored higher than that. For reference, here’s GPT‑Image‑2, NB Pro, and Muse compared: https://genai-showdown.specr.net/?models=nbp,g2,mi https://genai-showdown.specr.net/?models=nbp,g2,mi
- nrightnour 1mo agoThe photo rankings on this page are so absurd that the only reasonable explanation is that they were judged by an AI.
- minimaxir 1mo agoThe rankings are for prompt adherence, not subjective quality.