8 ms·
The Growing Compute Shortage
- oezi 2mo agoAt least one graph showing global total datacenter compute over time and projected to come online would have been helpful.
- no-name-here 2mo agohttps://epoch.ai/data-insights/ai-chip-production https://epoch.ai/data-insights/ai-chip-production from the beginning of 2026 claimed global AI compute capacity was doubling every 7 months.
- incognito124 2mo ago> from the beginning of 2026 claimed global AI compute capacity was doubling every 7 months. That's definitely one way to say "it doubled this year"
- computerphage 2mo agoI read it as "was posted earlier this year. It claimed that compute has been doubling every seven months for quite some time, on average"
- no-name-here 2mo agoNo, it was "<linked article> from the beginning of 2026 claimed <x>." Also, if you open the article, you'll see the data goes back multiple years, but does not cover 2026 (as it was posted at the beginning of 2026).
- Mistletoe 2mo agoGraphs like this are all I can think of when I see unsustainable euphoria like above. https://www.bigtrends.com/education/lessons-from-the-past-10-charts-graphs-of-the-great-depression/ https://www.bigtrends.com/education/lessons-from-the-past-10...
- jcattle 2mo agoAh I thought you would link https://xkcd.com/605/ https://xkcd.com/605/
- post-it 2mo agoNot what it says. > We fit a trend to the quarterly quantities of AI compute sold, as measured in H100-equivalents. We find that computing capacity has been growing by 3.3x per year, equivalent to a doubling time of 7 months. This is just a graph of cash changing hands. I transferred a loonie from one hand to the other one trillion times this morning and have the largest data centre in the world.
- p808 2mo ago[dead]
- twoodfin 2mo agoIf instead of us training the models to be human utility maximizers, the models had figured out how to train us to maximize their own utility, could we tell the difference?
- GolfPopper 2mo agoI suspect that utility may not enter the picture at all, and what we have are LLMs that have been optimized for getting the sort of humans who decide what to spend money on to maximize spending on LLMs and related infrastructure. The LLMs in turn could be anything from Skynet to a very spicy cold-reading autocomplete.
- Normal_gaussian 2mo agoWhat on earth is that H100 price trend graph. The equally spaced x-axis points are 2x6 months, 6x3 months, 9x1 month. The whole visualisation of the trend is ruined on the back of that. There are liars, damned liars, and people who play silly buggers with scales.
- nh23423fefe 2mo agoi dont see how compressing the past ruins the present day extrapolation? How does the conclusion change for you if the graph was 3x wider on the left? The title is "rates are rising" is that not true?
- danlitt 2mo agoThat specific conclusion is unaffected, but that doesn't make it okay. If they used fake numbers, but the conclusion were still the same, you presumably wouldn't think that was okay either?
- Normal_gaussian 2mo agoA graph is much more than one conclusion; in fact, almost the entire point of graphing is to allow the comparison of "shapes" and to easily hypothesise about associations across datasets. This graph misrepresents the rate the price declines and the length of time it has been stable for, which throws off nearly all non-trivial conclusions.
- nh23423fefe 2mo agoIt doesn't mis-represent it. If you think that, then you think log plots misrepresent rates.
- Normal_gaussian 2mo agoA) a log scale is self consistent B) be careful about telling other people what they think
- 2mo ago
- s08148692 2mo agoIt will be fascinating to watch this play out. Particularly if AI-compute satellite constellations become a reality - We're trending towards a matrioshka brain and I'll be happy if I live to see the beginnings of that future
- post-it 2mo agoThey'll become a reality on the same day as solar roadways.
- sneak 2mo agoGPUs aren’t scarce. There is just a lot of demand, so prices have gone up. Anyone can get GPUs right now and build out what they need; it’s just a matter of paying more for them than the next company. I’ve been looking at building a multi-TB HBM system lately and they’re readily available - they just cost $400k.
- gspr 2mo ago> GPUs aren’t scarce. There is just a lot of demand, so prices have gone up. What does "scarce" mean to you? I'd wager that to most reasonable people it means "available in a supply that, measured against demand, is low". Market forces typically react to such a state by pushing prices up. Saying "they're not scarce, they're just expensive" is just silly. The exception to this would be during e.g. supply chain hiccups (lots of compute sitting there, unable to get to its operational destination) or when other forces artificially push up prices. But that's not what we're seeing here. Compute is much scarcer than it's been for years.
- fhdkweig 2mo agoThere is the type of scarce when you can't get something for love or money. Look at the gasoline lines in Russia. Those people sit in lines miles long for gas stations that aren't even open. No matter how much money they are willing to spend, they still can't get the gas. I'm just grateful that I don't need to upgrade my computer for a while, and cross my fingers that I don't have a hardware failure in the next couple of years. I had an unpleasant laptop failure in 2021 that I don't wish to repeat.
- alt227 2mo agoThe definition of scarce is rare, or insufficient to meet demand. Not just that supply is low in relation to demand, it means that you physically have trouble getting something.
- gspr 2mo agoFine, replace "low" with "too low".
- RunSet 2mo ago> In 2026, compute is becoming the spice of our era. I think it more resembles the "content" of our era. https://refactoringenglish.com/blog/why-i-stopped-creating-content/ https://refactoringenglish.com/blog/why-i-stopped-creating-c...
- reticulates 2mo agoYes but how much of that compute shortage is from demand that is subsidized? We’ve seen companies like Uber drastically cut how much they are willing to spend on AI because they are paying actual usage costs, while at the same time OpenAI and Anthropic increase the limits on their fixed cost plans for individuals meaning people not paying usage costs are using it more and more… doesn’t this show that the compute shortage is because OpenAI and Anthropic are paying for it, not their customers? And the moment OpenAI and Anthropic stop paying for it, demand will collapse.
- rayiner 2mo agoThe compute demand is not fake. Non-coding industries have barely begun to deploy this technology. In the legal sector, I’ve been a tech pessimist my entire career, because it was uniformly quite bad. I’ve spent the last few months demoing legal tools backed by frontier models, and we’re definitely going to buy one of them. They’re real and they work and they address a bunch of needs.
- reticulates 2mo agoYou’re talking across the issue. The demand is real because it is cheap. The demand is being generated by OpenAI and Anthropic selling inference below cost on fixed price plans. If everyone was paying the actual costs then demand would fall through the floor. The legal tools you’re looking at use barely any compute. They’re not driving the compute demand. You can validate this by asking how much they are spending on API usage. A company spending $100,000 per month on a frontier model via an API is the equivalent of… 10 or so OpenAI and Anthropic fixed price plan customers. Are these legal tools spending hundreds of millions per year on the frontier models?
- lifeisstillgood 2mo agoMeh. That’s like saying we have a parking shortage in cities. No we have more car journies than we need. We have societies designed by the default choice. If that choice is walkable streets, low cost electric buses, dense neighbourhoods (mostly I mean you can walk for miles along streets and parks and shops without crossing a car park) Then you get far less car use. People aren’t stupid, but living in downtown Houston means you have far less choice about driving everywhere than living in suburban Amsterdam It’s all choices We just are making bad ones mostly The choices now are would you like ads or more ads. Shall I waste compute seeing if the web page you clicked on can be summarised ? The simple answer to AI is to charge it at cost - not subsidised. Then the market will start to shake out. It might take the US stock market with it …
- sghiassy 2mo agoIsn’t the compute shortage temporary? I don’t believe in 5-10yrs we’ll be in a shortage anymore
- gobdovan 2mo agoWe're in such a bull market, even shortages grow.
- khurs 2mo agoThis is written by a financial company that is pouring billions into data centres. https://www.apollo.com/insights-news/pressreleases/2026/01/apollo-backs-5-4-billion-valor-and-xai-data-center-compute-infrastructure-transaction-with-3-5-billion-capital-solution-3214463 https://www.apollo.com/insights-news/pressreleases/2026/01/a... https://www.apollo.com/insights-news/pressreleases/2025/11/apollo-funds-complete-acquisition-of-stream-data-centers-3179224 https://www.apollo.com/insights-news/pressreleases/2025/11/a... etc
- antr 2mo agoSo Apollo believes in the thesis strongly enough to put billions of its own capital behind it. That sounds more like putting their money where their mouth is than a rebuttal.
- TehCorwiz 2mo agoOr they invested billions and it behooves them to generate justification for that investment at a time when everyone else is also building railroads, I mean data centers.
- antr 2mo agoThat's possible, but it's still an argument about Apollo's incentives rather than the thesis itself. By all means scrutinize anyone talking their book, but then show where the analysis is wrong. Otherwise we've moved from “Apollo is conflicted” to “Apollo must be wrong because it invested,” which is not much of an argument.
- wkjagt 2mo agoWhen gas prices went up in the 70s because of fuel shortage, smaller (more fuel efficient) cars became more popular to use less fuel to do the same thing. I wonder if the same will happen with compute, by making software more efficient, and do the same thing with less compute.
- sdsdssweew213 2mo agoI bet everyone will soon use a thin client with 4GB of RAM, and all the compute happens in the cloud of some American corporation you pay subscription fees to.
- OGWhales 2mo agoI can see that happening, though that sounds pretty bleak
- Ekaros 2mo agoAnd that will probably fail. I don't think on average companies are capable anymore to deliver software that can operate on only 4GB... Even if everything but presentation layer is cloud based...
- wkjagt 2mo agoYou're probably right. On my 7th gen i7 with 16 GB, which should theoretically be plenty to check my email, GMail's web interface is very slow.
- zehaeva 2mo agoThe only thing I really hope for if this continues for an extended period of time is everyone optimizing their programs to use a minimal amount of ram.
- maxerickson 2mo agoFor the companies doing massive build outs, there's already significant pressure to improve effectiveness and efficiency.
- isodev 2mo agoI feel this post is blind to many of the secondary side effects of this "shortage". The rapid increase in prices and delivery times is having deleterious effects on all things tech - everything from phones to smart appliances and all sorts of gadgets has moved into unreachable price levels. Just imagine, if the rumours are true and Apple's 'foldable' phone costs 2500€ or more - who is going to buys this? So much of Apple's ecosystem depends on enough users consuming services, buying apps and using their phones to facilitate digital interactions. I've been waiting for 2 months to get a new mac mini for our office lab, and the Studios have moved into "we can't really justify this expense" price range. So the framing that AI is inevitable or projections that put the world into a year or more of this "scarcity", where people can no longer afford non-entry level gadgets, are just naive IMO.
- user43928 2mo agoI for one am going to buy Apple's foldable at the expected price between $2000-$2500. As long as they sell an iPhone 18 at a more accessible price point, I do not see the issue.
- isodev 2mo agoThe issue is that this will be another vision pro. The issue is that every year, we take device and services price increase (which of course nobody can do anything about because Apple-Google is a cartel based in the US where consumer protection is not even in the dictionary but that's a different topic probably) and nobody keeps track of just how much these things cost for a "not an influencer / person who works in tech / friend of the president" household.
- user43928 2mo agoI think it is likely the foldable will sell well. While it will be at a premium price point, people use their phones for hours every day, while the Vision Pro is a niche gadget. With videos, photos, web browsing, and reading being much better on a foldable, I can see this have mass appeal.
- 2mo ago
- cmiles8 2mo agoFour things are currently correct: 1. There is a huge demand for compute, specifically GPU compute 2. Infrastructure providers are building like crazy, including taking on massive debt to fund this because their own cash flow can’t cover the bills 3. The demand for that compute is broadly being paid for with investor dollars pumping up the valuation of AI companies, not cash flow from said companies. If those subsidies go away these companies can’t pay for the compute they’re buying. 4. Those that own a lot of compute are starting to offload it, looking for interested buyers (e.g., Meta looking to build a cloud biz or SpaceX selling its excess compute to others). All while advances in open weight models are making it appear that the major labs truly have no model moat. Put together those 4 things paint a very ugly business and financial picture that seems unlikely to just correct itself naturally. History tells us, very clearly, that “the way out” of such a scenario is a series of events that is likely to leave some of the current players severely damaged if not simply out of business.
- senordevnyc 2mo agoEh, I don’t think there’s any reason to think #3 is true, and the whole thing being a house of cards is predicated on that one. If the massive demand is still present for compute at market rates (which I believe it is), then your second point is just investors spending cap ex to build out valuable and profitable assets, no problem there. Time will tell.
- cmiles8 2mo agoWhat evidence says otherwise? OpenAI is projected to have massive losses for years to come and all indications are that’s driven by the cost of compute being far higher than the revenue generated by said compute.
- user43928 2mo agoAbout 4: Does it not make sense to rent out your compute if competitors have a better model and demand at higher prices? About the moat, Mythos was first made available to customers at the beginning of April. Kimi K3 is still behind this. Both OpenAI and Anthropic are expected to deploy significant upgrades in August.
- zer00eyz 2mo ago> Compute Is Becoming a Competitive Moat No, it is not a moat. The story here changes dramatically when you start to look at the facets of the industry that the poster is skipping over. > At the semiconductor level, TSMC’s advanced-node capacity—particularly N3, which underpins much of the AI accelerator ecosystem—is approaching full utilization through at least 2027. MS bought more GPU's than they had rack space for: https://www.datacenterdynamics.com/en/news/microsoft-has-ai-gpus-sitting-in-inventory-because-it-lacks-the-power-necessary-to-install-them/ https://www.datacenterdynamics.com/en/news/microsoft-has-ai-... Open AI bought out the memory: https://x.com/kwharrison13/status/2029248559388746168 https://x.com/kwharrison13/status/2029248559388746168 but they have no means to consume anywhere close to their order. Meanwhile both google and amazon are consuming a bunch of TSMC capacity to build their own ai chips, bypassing NVIDIA ... And as for them, they seem to be addicted to burning power to keep scaling, and that is a massive problem - if the next gen chips burn more watts for the same amount of work that is only going to exacerbate the power issues were having not help them. Tokens are just Gacha for business. https://en.wikipedia.org/wiki/Gacha_game https://en.wikipedia.org/wiki/Gacha_game - It is software you dont control and you are going to pay for a cache miss. That isnt a model that is sustainable (even more so in the authors multi agent flows). At the point that prices come down, (and they will) you're going to see a lot of corporations move from the cloud to on premise or back into colocation.
- kittikitti 2mo agoTo preface, this article presents valid points that I agree with. I would have also liked to know their findings on computational memory (CXL) and not just HBM or DDR. Also, a theoretical background on the Von Neumann and Harvard architectures would be helpful. Many people designing these systems understand that there are vast shortages in almost all sectors in computing hardware. I'm more interested in AI strategy of a future where there's an oversupply. I speculate that by 2030 there will be both an overproduction in the factories that make these chips and a burgeoning second hand market of AI GPU's. The differences in the compute requirements for training versus inference of AI will explain the hindsight in oversupply.
- drooby 2mo agoThis article is about silicon. But the other shortage that matters is meat-compute. AI is trained on human intelligence. The hyper-scalers are squeezing every last drop of automatically verified reward, and that may get us very far. A compiler passes or fails in milliseconds for free, forever. But.. "good design taste" has no compiler. Of course, taste isn't unverifiable. But it's is expensively verifiable. Noisy, slow, and orders of magnitude lower throughput. People with deep domain knowledge often can't articulate well _why_ one design works and the other doesn't. So, judgment arrives as a verdict, and not a crisp rationale. I guess we'll see if sample efficiency outpaces the cost of human judgement. In the mean time, leverage will sit with whoever holds this tacit knowledge (incumbents). I.e., hospital systems, law firms, chip designers, studios, SaaS that are dominating their niche.. and not with the labs training on it. To me, this is why valuations of companies like Palantir could potentially make sense.
- ACCount37 2mo ago"A compiler" isn't the bottleneck. The bottleneck of RLVR is the rollout. Before you even get the code for a compiler to do its thing, you need an AI to ingest and emit thousands of tokens autoregressively. That's what makes RLVR so expensive. Likewise - I don't think your "design taste" argument holds? Areas where "taste" exists are far easier for AI to tackle than areas where no data exists. It's easier to teach an AI how to make an image that looks good than to teach an AI how to de-solder a BGA chip. "Tacit knowledge" yes, "taste" - not really?
- redwood 2mo agoWhile building Fabs takes time you can only imagine that the amount of capital flooding in means there will be an explosion in infrastructure supply over the next decade and eventually prices will crater