6 ms·
Here is a trend I'm noticing: - GPT-5 mini costs $0.25/$2 and will be discontinued in December. - GPT-5.4 mini costs $0.75/$4.5 and is supposed to be the repl
by HyperL0gi 3mo ago
Here is a trend I'm noticing:
- GPT-5 mini costs $0.25/$2 and will be discontinued in December.
- GPT-5.4 mini costs $0.75/$4.5 and is supposed to be the replacement.
- GPT-5.4 nano costs $0.2/$1.25 and, while it ranks better in benchmarks than GPT-5 mini, it's not even close when you test it in real scenarios.
So you're left being forced to go to GPT 5.4 mini if you use 5 mini today.
The same thing is happening here as their “Luna“ model will cost $1/$6.
Can't we just stay with the models we actually want? I don't need GPT 5.4 mini. GPT-5 does the job.
Maybe it’s the realization that it was never that cheap in the first place and they're forcing us to upgrade in a slow and painful way.
- paxys 3mo agoIt’s the same as the SaaS model. Price keeps going up, and to justify it they keep forcing you to upgrade to new versions with features that nobody asked for.
- theptip 3mo ago“More intelligence” is the new feature. Almost everyone is asking for this. Citation: have you looked at OAI and Anthropic’s customer growth numbers?
- paxys 3mo agoEvery use case of every customer doesn’t need more intelligence. I’m willing to bet that the vast majority will be perfectly fine running on “low intelligence” at a cheap price forever.
- theptip 3mo agoI for sure agree that plenty of current use-cases are solvable by non-frontier models. However, you said “new versions with features that nobody asked for”, and I would prefer that you concede the point before shifting to arguing a new point. What customers are asking for is smarter models. Because the tasks that only smarter models can solve are higher value, higher margin, than the tasks that non-frontier models can solve.
- kolinko 3mo agoWhat are you talking about? Prices of lowest tiers of models have fallen how much - 10-100x over the last two years. And actually, the model quality you needed to pay for in the past, you can just run on device now essentially for free.
- gonzalohm 3mo agoYeah, this is the classic silicon valley strategy of selling at a loss and then once they have captured the market inflate prices. See Uber, Netflix, etc.
- simianwords 3mo agoThis is a constantly repeated conspiracy theory and is not true at all. The api costs do increase but aggregate costs per task decrease. The question is: do people need lower intelligence models at all? The answer is a resounding NO! How many people do you see using haiku or sonnet? I see very few and most people default to the latest model and just play with thinking effort. I think three layers are good enough and supporting more is not a good UX.
- gonzalohm 3mo agoDo I need the most intelligent model to generate boilerplate code, which is my main usage for AI? Resounding No. For my use case a model from a year ago is good enough
- unknownfuture 3mo agoI... use them all the time: plan with a more advanced model, build with a cheaper one. Anthropic literally packages a metamodel (opusplan) for that pattern. Also: calling the SV blitzscaling strategy of using VC money to fund loss leader products with the goal of building a monopoly via dumping a conspiracy is quite the position given there's entire books written in the topic...
- phainopepla2 3mo agoAre you only considering coding use cases? Many enterprise use cases, such as simple data extraction, are well served by cheaper models.
- CraigRood 3mo agoI don't see them capturing anything at this point. If inference was profitable then they could compete on price/model and capture the market. Then increase price and pay back the model training. Feels like they are just pulling in as much as they can whilst competing on capabilities instead. At which point its a case of who can last the longest. Doesn't feel like Uber/Netflix.
- neosat 3mo agoGood observations. There's definitely a trend in pricing increasing but also balanced by innovations and availability of other models (both open and closed) emerging as alternatives. It's natural for the labs to explore how much they can push pricing, and for competitors to explore how they can treat that margin as their opportunity to grow their business. Eventually the pricing should be more stable.
- benterix 3mo ago> Eventually the pricing should be more stable. Why do you think so? This game can be played forever, you just need strong marketing and orgs gullible enough to pay a higher price for a minor upgrade.
- malnourish 3mo agoHardware hosting old models isn't hosting new models. If you want consistent models, host your own open weights ones.
- wolttam 3mo agoIf you have no need for Anthropic/OpenAI's frontier model capability, you may be better served with an open-weight model that can't be taken away. Edit: > GPT-5 does the job. I bring up DeepSeek V4 Flash a lot on HN, but I want to mention that according to Artificial Analysis, it trades blows with GPT-5 (high) (from August, 2025) [0] [0]: https://artificialanalysis.ai/models/comparisons/deepseek-v4-flash-vs-gpt-5 https://artificialanalysis.ai/models/comparisons/deepseek-v4...
- paxys 3mo agoUnless you are hosting it yourself on your own infrastructure it absolutely can be taken away.
- amunozo 3mo agoBut you have multiple providers, not just one.
- cyanydeez 3mo agoI'm sure he's referring to the tightening of internet controls around social media as an extrapolation to controlling websites, etc.
- logicchains 3mo agoEven in that case it can't be taken away; GPT and Claude are banned in China yet there's still a huge black market for tokens.
- paxys 3mo agoAnd every single one of those providers would buckle under government pressure. Fable itself is hosted on all major cloud providers. How many offer it today?
- minimaxir 3mo ago
- simonw 3mo agoOn Nano "it's not even close when you test it in real scenarios" - what have you seen? What kind of things can GPT-5 Mini handle that GPT-5.4 Nano cannot?
- isamu_2000 3mo agoWe’re using GPT-5-mini in an enterprise data-processing workflow, and we too see that GPT-5.4 nano performs materially worse for our requirements, roughly 30% worse as measured through our test suite.
- barrell 3mo agoAlso can confirm gpt-5.4-nano was unable to even keep up with 4.1-mini. Had to move off of OpenAI once 4.1-mini was retired
- deleted 3mo ago[deleted]
- cyanydeez 3mo agoNo, you can't. These companies have two infrastructures: model training and model inference. Inference needs to cache, it can't cache random model data, so it's essentially dedicated; it can't spin up models on demand, it has to know what demand is coming. These companies are going to end up with very few models offered and that's probably generous. They might end up with just one model and you pay for removing it's safe guards.
- sourcecodeplz 3mo agowho tf would use mini when you have dsv4 flash
- tosh 3mo agodiscontinuing the cheaper options is a risky move for openai will trigger re-evaluations of models by other labs + inference providers
- HyperL0gi 3mo agoI can speak for myself. We are exactly at this moment trying to replace GPT 5 mini with an open weight / open source model. No luck so far.
- mistic92 3mo agoIts happening to Anthropic Haiku and Gemini Flash/Flash lite. All of them are increasing prices and deprecating cheap models.
- deleted 3mo ago[deleted]
- mchusma 3mo agoI've struggled with this. You definitely can have great cheap models. There are many of them open source and served profitably by neo-clouds. The big labs have basically given up on cheap models, and it is frustrating. It means applications are not likely to build as much on them anymore (we are shifting workloads from Haiku/Sonnet to Deepseek v4, for example). I suspect the problem is that they need to charge a lot to keep revenue numbers up, and they are more worried about cannibalizing themselves than others cannibalizing them.
- mips_avatar 3mo agoI think it's more that they're abandoning simpler AI tasks to chinese models. Qwen 35b and deepseek flash are better than gp5 mini on my tasks and way cheaper.
- theptip 3mo ago> Maybe it’s the realization that it was never that cheap in the first place and they're forcing us to upgrade in a slow and painful way. All the analysis I have seen points to frontier models being profitable to serve. It’s using 50% or more of your GPUs for research plus CapEx for capacity expansion that makes these businesses so heavily cash-negative. What you are observing is downstream of another detail. It gets more expensive to serve a model as utilization goes down. Plus the opportunity cost vs newer, more-profitable models. There are plenty of valid reasons to critique here. “OpenAI is lying about this being a sustainable price to serve” is not one of them.
- btbuildem 3mo ago> stay with the models we actually want If you want control over the models you use, you have to self-host.
- hadlock 3mo agoEach model release gives an opportunity to reduce the number of old models still on offer, and charge a higher, less-subsidized tier. The trick is to charge a subsidized price that is less than an M3 Ultra, so they continue paying you rent, instead of a one-time fixed cost. So far open models can't compete with Opus 4.5 but as soon as it can, people will be looking at buying devices that can run that model locally. We are a claude shop but we already bought two mac studios to start migrating less complex but still agentic workflows there. We will break even on those in less than a year.
- fouc 3mo agoBreaking even in less than a year? What's the math on that?
- CSMastermind 3mo ago5.5 is smart enough for 99% of my tasks. I need that level of intelligence at ever decreasing prices.
- aleksandrm 3mo agoI don't know about Cursor or other outlets, but I use GPT 5.4 exclusively in Windsurf (Sorry, Devin!), and it's a very capable model that doesn't break the bank!.
- abc123abc123 3mo agoNo. Welcome to the wonderful world of SaaS. If you want your gui, your terms, your software, self-host. But I think, in time, a new generation will relearn this truth.
- zeryx 3mo agoWhy not self host or go to openrouter if you don't need SOTA frontier?
- jeffybefffy519 3mo agoGemini has done the same thing, gemini-3.5-flash is 15x more expensive for input tokens than gemini-2.0-flash. They are forcing us up the pricing ladder by deprecating the old models....