9 ms·
GPT Image 1.5
https://platform.openai.com/docs/models/gpt-image-1.5 https://platform.openai.com/docs/models/gpt-image-1.5
- adammarples 9mo agoStill can't pass my image test Two women walking in single file Although it tried very hard and had them staggered slightly
- weird-eye-issue 9mo agoInterestingly when you Google that literally all of the images have two women walking side by side
- raw_anon_1111 9mo agoI still can’t get it to draw a “13 hour clock” correctly
- fellowniusmonk 9mo agoAll the latest round of openai is massively overfit.
- ChrisArchitect 9mo agoPost: https://openai.com/index/new-chatgpt-images-is-here/ https://openai.com/index/new-chatgpt-images-is-here/ (https://news.ycombinator.com/item?id=46291827 https://news.ycombinator.com/item?id=46291827)
- dang 9mo agoWe'll merge that thread hither to give some other submitters a chance.
- abbycurtis33 9mo agoI still use Midjourney, because all of these major players are so bad at stylistic and creative work. They're singularly focused on photorealism.
- xnx 9mo agoThis is surprising. Is there a gallery of images that illustrates this?
- takoid 9mo agoMidjourney has a gallery on their website: https://www.midjourney.com/explore https://www.midjourney.com/explore
- throwthrowuknow 9mo agotheir explore page is a firehose of examples created by users and you can see the prompt used so you can compare the results in other services https://www.midjourney.com/explore?tab=video_top https://www.midjourney.com/explore?tab=video_top
- FergusArgyll 9mo agoThat's the opinionated vs user choice dynamic. When the opinions are good, they have a leg up
- kingkawn 9mo agoThis is a cultural flaw that predates image generation. Even PG has made statements on HN in the past equating “rendering skill” with the quality of art works. It’s a stand-in for the much more difficult task of understanding the work and value of culture making within the context of the society producing it.
- doctorpangloss 9mo agoSuppose the deck for Midjourney hit Paul Graham's desk, and the CEO was just an average Y Combinator CEO - so no previous success story. He would have never invested in Midjourney at seed stage (meaning before launch / before there were users) even if he were given the opportunity. Better to read that particular story in the context of, "It would be very difficult to make a seed fund that is an index of all avant garde culture making because [whatever]."
- minimaxir 9mo agoI have a Nano Banana Pro blog post in the works expanding on my experiments with Nano Banana (https://news.ycombinator.com/item?id=45917875 https://news.ycombinator.com/item?id=45917875). Running a few of my test cases from that post and the upcoming blog post through this new ChatGPT Image model, this new model is better than Nano Banana but MUCH worse than Nano Banana Pro which now nails the test cases that previously showed issues. The pricing is unclear but gpt-image-1.5 appears to be 20% cheaper than the current gpt-image-1 model, which would put a `high`-quality generation in the same price range as Nano Banana Pro. One curious case demoed here in the docs is the grid use case. Nano Banana Pro can also generate grids, but for NBP grid adherence to the prompt collapses after going higher than 4x4 (there's only a finite amount of output tokens to correspond to each subimage), so I'm curious that OpenAI started with a 6x6 case albeit the test prompt is not that nuanced.
- vunderba 9mo agoI'll be running gpt-image-1.5 through my GenAI Showdown later today, but in the meantime if you want to see some legitimately impressive NB Pro outputs, check out: https://mordenstar.com/blog/edits-with-nanobanana https://mordenstar.com/blog/edits-with-nanobanana In particular, NB Pro successfully assembled a jigsaw puzzle it had never seen before, generated semi-accurate 3D topographical extrapolations, and even swapped a window out for a mirror.
- niklassheth 9mo agoNice! Your comparison site is probably the best one out there for image models
- jngiam1 9mo agoThe mirror test is cool!
- IgorPartola 9mo agoSubtle detail but the little table casts a shadow because of the light in the window and the shadow remains unchanged after the mirror replaces the window.
- neom 9mo agoAnyone else have issues verifying with openai? I always get a "congrats you're done" screen with a green checkmark from Persona, nothing to click, and my account stays unverified. (Edit, mystically, it's fixed..!)
- ezero 9mo agoEven from their own curated examples, this looks quite a bit worse than nano banan in terms of preserving consistency on image edits.
- almosthere 9mo agoI didn't have a good experience with NB. I am half Indian. Immediately changes my face to a prototypical Indian man every time I use it. This tool is keeping my look the same.
- gundmc 9mo agoI find including "don't change anything else" in the NBP prompt goes a long way.
- almosthere 9mo agoI tried all of those types of prompts
- mortenjorck 9mo agoNano Banana became useless for image edits once the safety training started rejecting anything as “I can’t edit some public figures.” My own profile picture? Can’t edit some public figures. A famous Norman Rockwell painting from 80 years ago? Can’t edit some public figures. Safety’d into oblivion.
- sharkjacobs 9mo agoWas it ever explained or understood why ChatGPT Images always has (had?) that yellow cast?
- kingkawn 9mo agoColloquially called the urine filter
- jebronie 9mo agolets not mince words, its called the "piss filter"
- minimaxir 9mo agoMy pet theory is that OpenAI screwed up the image normalization calculation and was stuck with the mistake since that's something that can't be worked around. At the least, it's not present in these new images.
- BoorishBears 9mo agoThere's still something off in the grading, and I suspect they worked around it (although I get what you mean, not easily since you already trained) I'm guessing when they get a clean slate we'll have Image 2 instead of 1.5. In LMArena it was immediately apparent it was an OpenAI model based on visuals.
- swyx 9mo agowdym it cant be worked around when there exist literal yellow tint corrector models/tools haha
- ineedasername 9mo agoYeah, though I can imagine a conversation like this: SWE: "Seriously? import PIL \ read file \ == (c + 10%, m = m, y = y, k = k) \ save file done!" Exec: "Yeah, and first blogger get's a hold of image #1 they generate, starts saying 'Hey! This thing's been color corrected w/o AI! lol lame'" Or not, no idea. i've not understood the choice either, besides very intelligent AI-driven auto-touch up for lighting/color correction has been a thing for a while. It's just, for those I end up finding an answer for, maybe 25% of head scratcher decisions do end of having a reasonable, if non intuitive answer for. Here? haven't been able to figure one yet though, or find a reason/mention by someone who appears to have an inside line on it.
- deleted 9mo ago[deleted]
- xnx 9mo agoGreat to have continued competition in the different model types. What angle is there for second tier models? Could the future for OpenAI be providing a cheaper option when you don't need the best? It seems like that segment would also be dominated by the leading models. I would imagine the future shakes out as: first class hosted models, hosted uncensored models, local models.
- pdevr 9mo ago>Now remove the two men, just keep the dog, and put them in an OpenAI livestream that looks like the attached image. Where is the image given along with the prompt? If I didn't miss it: Would have been nice to show the attached image.
- taytus 9mo agoon top of the prompt. It has a weird layout; I had to scroll up to see it.
- pdevr 9mo agoWhat confused me was this: "Now remove the two men, just keep the dog, and put them in an OpenAI livestream that looks like the attached image.". The word "them" implied plural. I was looking for the dog and something else. Thanks.
- catigula 9mo agoNano Banana Pro is so good that any other attempt feels 1-2 generations behind.
- Jonovono 9mo agoNano banana pro is almost as good as seedream 4.5!
- BoorishBears 9mo agoSeedream 4.5 is almost as good as Seedream 4! (Realistically, Seedream 4 is the best at aesthetically pleasing generation, Nano Banana Pro is the best at realism and editing, and Seedream 4.5 is a very strong middleground between the two with great pricing) gpt-image-1.5 feels like OpenAI doing the bare minimum to keep people from switching to Gemini every time they want an image.
- mohsen1 9mo agoUnlike Nano Banana it allows generating photos of children. Always fun to ask AI to imagine children of a couple but it's also kinda concerning that there might be terrible use cases.
- r053bud 9mo agoI was able to generate photos of my imagined children via Nano Banana
- hexage1814 9mo agoIf memory serves me, Nano Banana allows generating/editing photos of children. But anything that could be misinterpreted, gets blocked, even absolutely benign and innocent things (especially if you are asking to modify a photo that you upload there). So they allow, but they turn on the guardrails to a point that might not be useful in many situations.
- BoorishBears 9mo agoI haven't seen that, meanwhile gpt-image-1.5 still has zero-tolerance policing copyright (even via the API) so it's pretty much useless in production once exposed to consumers. I'm honestly surprised they're still on this post-Sora 2: let the consumer of the API determine their risk appetite. If a copyright holder comes knocking, "the API did it" isn't going to be a defense either way.
- rvz 9mo agoAnother bunch of "startups" have been eliminated.
- moralestapia 9mo agoAmong those, Photoshop.
- koakuma-chan 9mo agoI wish. Even Nano Banana Pro still sucks for even basic operations.
- alasano 9mo agoIt's still not available in the API despite them announcing the availability. They even linked to their Image Playground where it's also not available.. I updated my local playground to support it and I'm just handling the 404 on the model gracefully https://github.com/alasano/gpt-image-1-playground https://github.com/alasano/gpt-image-1-playground
- minimaxir 9mo agoIt's a staggered rollout but I am not seeing it on the backend either.
- joshstrange 9mo ago> staggered rollout It's too bad no OpenAI Engineers (or Marketers?) know that term exists. /s I do not understand why it's so hard for them to just tell the truth. So many announcements "Available today for Plus/Pro/etc" really means "Sometime this week at best, maybe multiple weeks". I'm not asking for them to roll out faster, just communicate better.
- anonfunction 9mo agoYeah I just tried it and got a 500 server error with no details as to why: POST "https://api.openai.com/v1/responses": 500 Internal Server Error { "message": "An error occurred while processing your request. You can retry your request, or contact us through our help center at help.openai.com if the error persists. Please include the request ID req_******************* in your message.", "type": "server_error", "param": null, "code": "server_error" } Interestingly if you change to request the model foobar you get an error showing this: POST "https://api.openai.com/v1/responses": 400 Bad Request { "message": "Invalid value: 'blah'. Supported values are: 'gpt-image-1' and 'gpt-image-1-mini'.", "type": "invalid_request_error", "param": "tools[0].model", "code": "invalid_value" }
- weird-eye-issue 9mo agoMy Enterprise account got an email 1.5 hours ago that it is available in API but my other accounts haven't gotten any email yet
- 0dayman 9mo agonah Nano Banana Pro is much better
- kitsune1 9mo ago[dead]
- StarterPro 9mo agoIn the image they showed for the new one, the mechanic was checking a dipstick...that was still in the vehicle. I really hope everyone is starting to get disillusioned with OpenAI. They're just charging you more and more for what? Shitty images that are easy to sniff out? In that case, I have a startup for you to invest in. Its a bridge-selling app.
- czhu12 9mo agoHaven’t their prices stayed at $20/m for a while now?
- wahnfrieden 9mo agoThey've published anticipated price increases over coming years. Prices will rise dramatically and steadily to meet revenue targets.
- cheema33 9mo agoAI doesn’t have much of a moat. People can and will easily switch providers.
- wahnfrieden 9mo agoSure but there are only a couple leading providers worth considering for coding at least, and there will be consolidation once investment pulls back. They may find a way to collude on raising prices. Where switching will be easier is with casual chat users plus API consumers that are already using substandard models for cost efficiency. But there will also always be a market for state of art quality.
- wahnfrieden 9mo agoReinforced today: As Gemini has gained competitiveness (higher confidence in its output, better reputation), its prices have steadily risen
- blurbleblurble 9mo agoIt's really weird to see "make images from memories that aren't real" as a product pitch
- nurettin 9mo agoIt would creep me out if the model produced origami animals for that prompt.
- kingstnap 9mo agoIt's strange to me too, but they must have done the market research for what people do with image gen. My own main use cases are entirely textual: Programming, Wiki, and Mathematics. I almost never use image generation for anything. However its objectively extremely popular. This has strong parallels for me to when snapchat filters became super popular. I know lots of people loved editing and filtering pictures but I always left everything as auto mode, in fact I'd turn off a lot of the default beauty filters. It just never appealed to me.
- 999900000999 9mo agoI can actually imagine actors selling the rights to make fake images with them. In late stage capitalism you pay for fake photos with someone. You have chat gpt write about how you dated for a summer, and have it end with them leaving for grad school to explain why you aren't together. Eventually we'll all just pay to live in the matrix. When your credit card is declined you'll be logged out, to awaken in a shared studio apartment. To eat your rations.
- ares623 9mo agoI can see them getting paid like residuals from TV re-runs. But after a point it'll hit saturation point. The novelty will wear off since everyone has access to it. Who cares if you have a fake photo with a celebrity if everyone knows it's fake.
- oblio 9mo ago> When your credit card is declined you'll be logged out, to awaken in a shared studio apartment. To eat your rations. You're funny. No, you'll awaken in a tent, next to your shopping cart, under the bridge.
- oxag3n 9mo agoIf this was a farm of sweatshop Photoshopers in 2010, who download all images from the internet and provide a service of combining them on your request, this would escalate pretty quickly. Question: with copyright and authorship dead wrt AI, how do I make (at least) new content protected? Anecdotal: I had a hobby of doing photos in quite rare style and lived in a place where you'd get quite a few pictures of. When I asked gpt to generate a picture of that are in that style, it returned highly modified, but recognizable copy of a photo I've published years ago.
- nobody_r_knows 9mo agomy question to your anecdotal: who cares? not being fecicious, but who cares if someone reproduced your stuff and millions of people see your stuff? is the money that you want? is it the fame? because fame you will get, maybe not money... but couldn't there be another way?
- netule 9mo agoSuddenly, copyright doesn't matter anymore when it's no longer useful to the narrative.
- BoorishBears 9mo agoOpenAI does care about copyright, thankfully China does not: https://imgur.com/a/RKxYIyi https://imgur.com/a/RKxYIyi (to clarify, OpenAI stops refining the image if a classifier detects your image as potentially violating certain copyrights. Although the gulf in resolution is not caused by that.)
- ragequittah 9mo agoCopyright has overstepped its initial purpose by leaps and bounds because corporations make the law. If you're not cynical about how Copyright currently works you probably haven't been paying attention. And it doesn't take much to go from cynical to nihilist in this case.
- surrTurr 9mo agonot super impressed. feels like 70% as good as nano banana pro.
- sfmike 9mo agoHope to see more "red alert" status from the ai wars putting companies into al hands on deck. This is only helping cost of tokens and efficacy. As always competition only helps the end users.
- youknow123 9mo ago[dead]
- dzonga 9mo agowe seriously can't be burning GW of energy just to have sama in a GPT-Shirt Ad generated by A.I impressive stuff though - as you can give it a base image + prompt.
- drawnwren 9mo agocounterpoint: we should make energy abundant enough that it really doesn't matter if sama wants to generate gpt-shirt ads or not. we have the capability, we just stopped making power more abundant.
- iknowstuff 9mo agoI think we can say the pause we took was reasonable once we realized the environmental impact of dumping greenhouse gases into the atmosphere but if now that can ensure further growth won’t do it, let’s make sure we restart, just clean this time.
- astrange 9mo agoIt's a joke about one of his old fits. https://x.com/coldhealing/status/1747270233306644560 https://x.com/coldhealing/status/1747270233306644560
- KaiserPro 9mo agoIs there a watermarking, or some other way for normal people to tell if its fake?
- PhilippGille 9mo agohttps://help.openai.com/en/articles/8912793-c2pa-in-chatgpt-images https://help.openai.com/en/articles/8912793-c2pa-in-chatgpt-... It doesn't mention the new model, but it's likely the same or similar.
- adrian17 9mo agoI just checked several of the files uploaded to the news post, the "previous" and "new", both the png and webp (&fm=webp in url) versions - none had the content metadata. So either the internal version they used to generate them skipped them, or they just stripped the metadata when uploading.
- mmh0000 9mo agoI know OpenAI watermarks their stuff. But I wish they wouldn't. It's a "false" trust. Now it means whoever has access to uncensored/non-watermarking models can pass off their faked images as real and claim, "Look! There's no watermark, of course, it's not fake!" Whereas, if none of the image models did watermarking, then people (should) inherently know nothing can be trusted by default.
- pbmonster 9mo agoYeah, I'd go the other way. Camera manufacturers should have the camera cryptographically sign the data from the sensor directly in hardware, and then provide an API to query if a signed image was taken on one of their cameras. Add an anonymizing scheme (blind signatures or group signatures), done.
- mnorris 9mo agoI ran exiftool on an image I just generated: $ exiftool chatgpt_image.png ... Actions Software Agent Name : GPT-4o Actions Digital Source Type : http://cv.iptc.org/newscodes/digitalsourcetype/trainedAlgorithmicMedia http://cv.iptc.org/newscodes/digitalsourcetype/trainedAlgori... Name : jumbf manifest Alg : sha256 Hash : (Binary data 32 bytes, use -b option to extract) Pad : (Binary data 8 bytes, use -b option to extract) Claim Generator Info Name : ChatGPT ...
- gostsamo 9mo agoAlt text is one of the nicest uses for ai and still Open AI didn't bother using it for something so basic. The dogfooding is not strong with their marketing team.
- zkmon 9mo agoAI-generated images would remove all the trust and admire for human talent in art, similar to how text-generation would remove trust and admire for human talent in writing. Same case for coding. So, let's simulate that future. Since no one trusts your talent in coding, art or writing, you wouldn't care to do any of these. But the economy is built on the products and services which get their value based how much of human talent and effort is required to produce them. So, the value of these services and products goes down as demand and trust goes down. No one knows or cares who is a good programmer in the team, who is great thinker and writer and who is a modern Picasso. So, the motivation disappears for humans. There are no achievements to target, there is no way to impress others with your talent. This should lead to uniform workforce without much difference in talents. Pretty much a robot army.
- arnz-arnz 9mo agoall I can hope for is that a new industry or reliable ecosystem of vetters of real human talent will emerge. Are you really as good a writer as you claim to be? Show us the badge. That or AI firms have to be forced to 'watermark' all their creative outputs, and anyone misleading the public/audience should be punishable by law.
- zkmon 9mo agoBoth are just mid-summer dreams. There is no global law to enforce watermark. There are no badges that can't be forged.
- arnz-arnz 9mo agoThere isn't but that doesn't mean there won't be. It can even go as far as banning certain features. There isn't just hope with the kind of politics we have right now.
- deleted 9mo ago[deleted]
- nycdatasci 9mo ago[flagged]
- aziis98 9mo agoI know this is a bit out of scope for these image editing models but I always try this experiment [1] of drawing a "random" triangle and then doing some geometric construction and they mess up in very funny ways. These models can't "see" very well. I think [2] is still very relevant. [1]: https://chatgpt.com/share/6941c96c-c160-8005-bea6-c809e58591c1 https://chatgpt.com/share/6941c96c-c160-8005-bea6-c809e58591... [2]: https://vlmsareblind.github.io/ https://vlmsareblind.github.io/
- anonfunction 9mo agoSo the announcement said the API works with the new model, so I updated my Golang SDK grail (https://github.com/montanaflynn/grail https://github.com/montanaflynn/grail) to use but it returns a 500 server error when you try to use it, and if you change to a completely unknown model it's not listed in the available models: POST "https://api.openai.com/v1/responses": 500 Internal Server Error { "message": "An error occurred while processing your request. You can retry your request, or contact us through our help center at help.openai.com if the error persists. Please include the request ID req_******************* in your message.", "type": "server_error", "param": null, "code": "server_error" } POST "https://api.openai.com/v1/responses": 400 Bad Request { "message": "Invalid value: 'blah'. Supported values are: 'gpt-image-1' and 'gpt-image-1-mini'.", "type": "invalid_request_error", "param": "tools[0].model", "code": "invalid_value" }
- gs17 9mo ago> Still some scientific inaccuracies, but ~70% correct That's still dangerously bad for the use-case they're proposing. We don't need better looking but completely wrong infographics.
- astrange 9mo agoIt's pretty common for infographics to be wrong. The people making them aren't the same people who know the facts. I'd especially say like 100% of amateur political infographics/memes are wrong. ("climate change is caused by 100 companies" for instance)
- rcarmo 9mo agoWe don’t, but most Marketing departments salivate for them.
- agentifysh 9mo agoI am very impressed a benchmark I like to run is have it create sprite maps, uv texture maps for an imagined 3d model Noticed it captured a megaman legends vibe .... https://x.com/AgentifySH/status/2001037332770615302 https://x.com/AgentifySH/status/2001037332770615302 and here it generated a texture map from a 3d character https://x.com/AgentifySH/status/2001038516067672390/photo/1 https://x.com/AgentifySH/status/2001038516067672390/photo/1 however im not sure if these are true uv maps that is accurate as i dont have the 3d models itself but ive tried this in nano banana when it first came out and it couldn't do it
- gs17 9mo ago> however im not sure if these are true uv maps I can tell you with 100% certainty they are not. For example, Crash doesn't have a backside for his torso. You could definitely make a model that uses these as textures, but you'd really have to force it and a lot of it would be stretched or look weird. If you want to go this approach, it would make a lot more sense to make a model, unwrap it, and use the wireframe UV map as input. Here's the original Crash model: https://models.spriters-resource.com/pc_computer/crashbandicootthensanetrilogy/asset/314827/ https://models.spriters-resource.com/pc_computer/crashbandic... , its actual texture is nothing like the generated one, because the real one was designed for efficiency.
- agentifysh 9mo agoyeah definitely impressive compared to what nano banana outputted tried your suggested approach by unwrapaped wireframe uv as input and im impressed https://x.com/AgentifySH/status/2001057153235222867 https://x.com/AgentifySH/status/2001057153235222867 obviously its not going to be accurate 1:1 but with more 3d spatial awareness i think it could definitely improve
- Nition 9mo agoThat's a remake model in a modern game. The original Crash was even simpler than that one. Most of Crash in the first game was not textured; just vertex colours. Only the fur on his back and his shoelaces were textures at all.
- brador 9mo agoEvery person in every picture in their examples is white except for 1 Asian dude. Like a 46:1 ratio for the page (I counted). Not one Middle Eastern or Black or Jewish or Indian or South American person. Not even one. And no one on the team said anything? Come on Sam, do better.
- gpt-image 9mo ago[dead]
- celeryd 9mo agoIf it can't generate non-sexual content of a woman in a bikini, I am not interested.
- ares623 9mo agoMy copium is that analog photography makes a come back as a way to recover some level of trust and authenticity.
- Forgeties79 9mo agoGood luck getting it developed unfortunately. I have to ship it off now, there isn’t a single local spot in my city that will develop anymore
- ares623 9mo agoWhen the demand is back, the labs should start coming back. There's a few in my relatively small city which is pretty surprising. But the costs are still too high to cover the low volume I guess.
- Forgeties79 9mo agoThe big issue is chemical disposal IIRC (which yes is a cost just being more specific)
- famahar 9mo agoI was reading a trend report on art and it seems like collage, squiggly hand drawn text, and lots of intentional imperfections are becoming popular. I'm not sure how hard it is for AI to recreate those, but it is nice to see people trying to do more of what AI struggles with.
- vunderba 9mo agoOkay results are in for GenAI Showdown with the new gpt-image 1.5 model for the editing portions of the site! https://genai-showdown.specr.net/image-editing https://genai-showdown.specr.net/image-editing Conclusions - OpenAI has always had some of the strongest prompt understanding alongside the weakest image fidelity. This update goes some way towards addressing this weakness. - It's leagues better at making localized edits without altering the entire image's aesthetic than gpt-image-1, doubling the previous score from 4/12 to 8/12 and the only model that legitimately passed the Giraffe prompt. - It's one of the most steerable models with a 90% compliance rate Updates to GenAI Showdown - Added outtakes sections to each model's detailed report in the Text-to-Image category, showcasing notable failures and unexpected behaviors. - New models have been added including REVE and Flux.2 Dev (a new locally hostable model). - Finally got around to implementing a weighted scoring mechanism which considers pass/fail, quality, and compliance for a more holistic model evaluation (click pass/fail icon to toggle between scoring methods). If you just want to compare gpt-image-1, gpt-image-1.5, and NB Pro at the same time: https://genai-showdown.specr.net/image-editing?models=o4,nbp,g15 https://genai-showdown.specr.net/image-editing?models=o4,nbp...
- echelon 9mo agoI really love everything you're doing! Personal request: could you also advocate for "image previz rendering", which I feel is an extremely compelling use case for these companies to develop. Basically any 2d/3d compositor that allows you to visually block out a scene, then rely on the model to precisely position the set, set pieces, and character poses. If we got this task onto benchmarks, the companies would absolutely start training their models to perform well at it. Here are some examples: gpt-image-1 absolutely excels at this, though you don't have much control over the style and aesthetic: https://imgur.com/gallery/previz-to-image-gpt-image-1-x8t1ijX https://imgur.com/gallery/previz-to-image-gpt-image-1-x8t1ij... Nano Banana (Pro) fails at this task: https://imgur.com/a/previz-to-image-nano-banana-pro-Q2B8psd https://imgur.com/a/previz-to-image-nano-banana-pro-Q2B8psd Flux Kontext, Qwen, etc. have mixed results. I'm going to re-run these under gpt-image-1.5 and report back. Edit: gpt-image-1.5 : https://imgur.com/a/previz-to-image-gpt-image-1-5-3fq042U https://imgur.com/a/previz-to-image-gpt-image-1-5-3fq042U And just as I finish this, Imgur deletes my original gpt-image-1 post. Old link (broken): https://imgur.com/a/previz-to-image-gpt-image-1-Jq5M2Mh https://imgur.com/a/previz-to-image-gpt-image-1-Jq5M2Mh Hopefully imgur doesn't break these. I'll have to start blogging and keep these somewhere I control.
- mingabunga 9mo agoDid an experiment to give a software product a dark theme. Gave Both (GPT and Gemini/Nano) a screenshot of the product and an example theme I found on Dribbble. - Gemini/Nano did a pretty average job, only applying some grey to some of the panels. I tried a few different examples and got similar output. - GPT did a great job and themed the whole app and made it look great. I think I'd still need a designer to finesse some things though.
- smlavine 9mo agoThis is terrifying. Truth is dead.
- WhyOhWhyQ 9mo agoMakes you wonder what's really meant when we talk about progress.
- teaearlgraycold 9mo agoEventually phone manufacturers will be forced to become arbiters of truth with signed images and videos.
- ge96 9mo agoI get the tech implementation is amazing, I wonder if it takes away from genuineness of events, like the Astronaut photo, I get it's just a joke/funny too but it's like a photo of you in a supercar vs. actually buying one. Or fake AI companions vs. real people. Beauty filters/skinny filters vs. actually being healthy.
- onoesworkacct 9mo agothe next generation of humans growing up will not even care whether media is real or not any more. The saturation of AI content and FUD around real content is going to blur the lines to the extent that there's no point even caring about it. And it's an intractable problem. hopefully this leads to greater importance of seeing things with your own wetware.
- ge96 9mo agoThe other issue is the need to show off... if I had a supercar why do I have to post it on Instagram that kind of thing.
- ge96 9mo agothinking about this more it goes two ways for guys/girls, the guys can post pics of them doing crazy things on their Tinder up to the girl to decide if it's real or not I'm not saying this as a critique against image generation as you can manually make these fake images but yeah Ultimately I think it's good, makes people be real
- enigma101 9mo agoReally can't stand the image slop suffocating the internet.
- eterm 9mo agoI have a "go to" prompt for images: > In the style of a 1970s book sci-fi novel cover: A spacer walks towards the frame. In the background his spaceship crashed on an icy remote planet. The sky behind is dark and full of stars. Nano banana pro via gemini did really well, although still way too detailed, and it then made a mess of different decades when I asked it to follow up: https://gemini.google.com/share/1902c11fd755 https://gemini.google.com/share/1902c11fd755 It's therefore really disappointing that GPT-image 1.5 did this: https://chatgpt.com/share/6941ed28-ed80-8000-b817-b174daa922a7 https://chatgpt.com/share/6941ed28-ed80-8000-b817-b174daa922... Completely generic, not at all like a book cover, it completely ignored that part of the prompt while it focused on the other elements. Did it get the other details right? Sure, maybe even better, but the important part it just ignored completely. And it's doing even worse when I try to get it to correct the mistake. It's just repeating the same thing with more "weathering".
- bongodongobob 9mo agoYou're just not describing what you want properly. Looks fine to me. Clearly you have something else in mind, so I think you're just not describing well. My tip would be to use actuall illustration language. Do you want a wide angle shot? What should depth of field be? Oil painting print? Ink illustration? What kind of printing style? Do you want a photo of the book or a pre-print proof? What kind of color scheme? A professional artist wouldn't know what you want. You didn't even specify an art style. 1970s sci-fi novel cover isn't a style. You'll find vastly different art styles from the 70s. If you're disappointed, it's because you're doing a shitty job describing what's in your head. If your prompt isn't at least a paragraph, you're going to just get random generic results.
- eterm 9mo agoThe killer feature of LLMs is to be able to extrapolate what's really wanted from short descriptions. Look again at Gemini's output, it looks like an actual book cover, it looks like an illustration that could be found on a book. It takes on board corrections (albeit hilariously literaly). Look at GPT image's output, it doesn't look anything like a book cover, and when prompted to say it got it wrong, just doubles down on what it was doing.
- nightshift1 9mo agoWhat is the endgame? Why is OpenAI throwing that much money on image/video generation? Is there a profitable market for AI-generated image slop? Do people choose ChatGPT instead of Gemini/Grok/Claude because of the image generation capabilities? To me, it looks like a huge fiery money pit.
- BrokenCogs 9mo agoThe endgame is to make money during the hype and then cash out before it crashes.
- bdangubic 9mo agoif that is the endgame openai is doing everything but working towards that goal :)
- BrokenCogs 9mo agoYeah they fumbled big time
- password-app 9mo agoImpressive image quality improvements. Meanwhile, AI agents just crossed a milestone: Simular's Agent S hit 72.6% on OSWorld (human-level is 72.36%). We're seeing AI get better at both creative tasks (images) and operational tasks (clicking through websites). For anyone building AI agents: the security model is still the hard part. Prompt injection remains unsolved even with dedicated security LLMs.
- randall 9mo agodouble popped collar ftw
- hamonrye 9mo ago[dead]
- encroach 9mo agoThis outperforms Gemini 3 pro image (nano banana pro) on Text-to-Image Arena and Image Edit Arena. I'm surprised they didn't mention this leaderboard in the blog post. I like this benchmark because its based upon user votes, so overfitting is not as easy (after all, if users prefer your result, you've won). https://lmarena.ai/leaderboard/text-to-image https://lmarena.ai/leaderboard/text-to-image https://lmarena.ai/leaderboard/image-edit https://lmarena.ai/leaderboard/image-edit
- nycdatasci 9mo agoThe arena concept doesn’t work for image models due to watermarks.
- encroach 9mo agoThere are no watermarks in the arena.
- nycdatasci 9mo agoThere are no visible watermarks, but model makers can use steganographic codes to identify outputs from their own models.
- nycdatasci 9mo agoText-to-Image Models Leave Identifiable Signatures: Implications for Leaderboard Security https://arxiv.org/pdf/2510.06525 https://arxiv.org/pdf/2510.06525
- encroach 9mo agoThis is true, however LMArena does employ some methods to mitigate attempts to manipulate the leaderboard, see https://openreview.net/forum?id=zf9zwCRKyP https://openreview.net/forum?id=zf9zwCRKyP They also control for style https://news.lmarena.ai/sentiment-control/ https://news.lmarena.ai/sentiment-control/
- GaryBluto 9mo agoGod OpenAI are so far behind. Their own example shows that trying to only change specific parts of the image doesn't work without affecting the background.
- thumbsup-_- 9mo agonow you can create good memories with your family without meeting them
- augustk 9mo agoOr create the family
- andai 9mo agoSam Altman Christmas decoration isn't real, he can't hurt me...
- sroussey 9mo ago“ Photo of a blond male in his 50s with half gray hair “ Still fails. Every photo of a man with half gray hair will have the other half black.
- v9v 9mo agoLots of em-dashes in this copy.
- jdthedisciple 9mo agoWhy is the emphasis of these promos always to create fake social media pictures of people and things that didnt happen? Aren't we plagued enough by all the fake bullshit out there. Ffs! /rant Sorry gotta be honest and blunt every one of those times...
- rw2 9mo agoHaving used it compared to Nano Banana: -The latency is still too high, lower than 10 seconds for nano banana and around 25 seconds for GPT image 1.5 -The quality is higher but not a jump like previous google models to Nano Banana Pro. Nano banana pro is still at least equivalently good or better in my opinion.
- chakintosh 9mo agoCan't wait to generate fake memories with my 20 years ago dead grandma
- bunnybomb2 9mo agoOr me and my ex
- Garlef 9mo agoGPT images is the new MS Word "Arial + clip art"
- sipsi 9mo agothe combination of two images the last gpt-image (nano banana) generated seem to be inappropriate
- animanoir 9mo ago[dead]
- fock 9mo agoGood to see that hands are still not solved...
- yuni_aigc 9mo agoOne thing I’ve noticed when comparing these models is that “quality” and “realism” don’t always move together. Some models are very strong at sharp details and localized edits, but they can break global lighting consistency — shadows, reflections, or overall scene illumination drift in subtle ways. GPT-Image seems to trade a bit of micro-detail for better global coherence, especially in lighting, which makes composites feel more believable even if they’re not pixel-perfect. It’s hard to capture this in benchmarks, but for real-world editing workflows it ends up mattering more than I initially expected.
- aymenfurter 9mo agoBig jump from gpt-image, but I would love to see the same kind of step change we have seen in recent language model releases.
- JustinXie 9mo ago[dead]
- JustinXie 9mo ago[dead]