7 ms·
In many investment theses - like Nvidia's bet that demand for compute will keep growing - the first order assumption is usually correct. Yes, demand for more co
by jcfrei 1mo ago
In many investment theses - like Nvidia's bet that demand for compute will keep growing - the first order assumption is usually correct. Yes, demand for more compute, chips, infrastructure is huge and each year some additional data centers will be built. Where such investment bets usually fail is in the second-order assumptions: Ie. the expectation of the growth of demand. This is where there's a high chance that the current expectations are likely exaggerated. So: demand is likely to persist for the foreseeable future but not increase every year. And that can upend the whole investment story. That can be enough to make these bonds a huge burden for Nvidia in the end. Not because people stopped buying more compute but because they stopped buying more every year.
- onlyrealcuzzo 1mo agoWhat makes this insanely hard to predict is that the compute needed for the same quality output has roughly gone down 90% every 18 months for ~5 years. 1) We don't know how long that trend will continue, but you do know where to look for when it may end (if smaller sized models continue to compress the knowledge effectively of larger models). 2) We don't know when the appetite for higher cost models might go down and by how much if smaller models get "good enough" and price becomes far more important. It is entirely possible that 5 years from now, there's >100x LLM inference going on - but demand for AI chips (including memory) is only 2x or less. It is also entirely possible that at some size - LLMs pick up some emergent capability that doesn't scale well to smaller sizes - and that there's an incredible boost to demand to get that capability. It's just very hard to predict.
- mattnewton 1mo agoI think efficiency is unlikely to result in lower demand for compute, instead more useful compute per watt increases the value of that compute; and we are not going to run out of economically useful things to do with it anytime soon on the demand side. The harder thing to forecast for me is if we hit a wall on increasing efficiency, either on the model weights side or silicon side, with current approaches. If we have to switch to something like burning the model weights into silicon to continue to make gains, then the current math on general purpose accelerators might be upside down.
- ekunazanu 1mo agoI agree; I don't think there's any reason to assume Jevons paradox won't apply. > If we have to switch to something like burning the model weights into silicon to continue to make gains I think that's already being considered semi-seriously [0][1] [0] https://taalas.com/products/ https://taalas.com/products/ [1] https://ir.amd.com/news-events/press-releases/detail/1296/amd-acquires-taalas-to-advance-compute-solutions-for-rapidly-growing-ai-inference-market https://ir.amd.com/news-events/press-releases/detail/1296/am...
- kurthr 1mo agoWhat's really interesting is that if you scale it to higher densities (eg 3nm and stacked die) with ComputeInMemory for fp8 you can reasonably start to fit 30B-70B models. With MoE and multiple stacked die, just like HBM, you could fit an open weight near frontier 1T model (like GLM5.2) at similarly much lower power <10kW and high token rates >2ktps. For running a bunch of agents where fill rate and speed/latency are important it may not matter that you're 6-18mo behind on weights. The process for the chip design could be largely automated, and new silicon pumped out as new weights are available (with a 3-6mo delay).
- adrianN 1mo agoWhen efficiency reaches the point where local models on consumer hardware are good enough, demand for cloud tokens could rapidly shrink.
- pseudosavant 1mo agoIt could be quite a while before we reach that point though. 5+ years easily. I've been keenly interested in the ability to run local models, but the hardware is just not there. Consumer RAM speeds and capacity will have to significantly increase before local models will be able to perform as well as even the lowest end GPT-5.6 Luna model. This is on the backdrop of RAM becoming prohibitively expensive. And without the speed and quantity of RAM, it becomes impossible to generate tokens at interactive speeds, regardless of model. There is a fundamental dependency between calculating all of the active params with the given RAM speed. Even with a model that has been quantized all the way down to Q4, the DGX/RTX Spark chip with 128GB of RAM can only generate ~18 tokens/sec for a MoE model with only 30B active parameters. There haven't been any broadly useful models below 30B active parameters. And that is for a $5000+ piece of hardware that will be one of the best for running on-device models. I really want to buy instead of rent my AI, but the economics are truly terrible.
- amelius 1mo agoNvidia's great superpower is flexibility. You can easily run models of very different types on the same card; and their hardware is great for R&D. However, at some point AI may be good enough for most people and then it makes sense to make an ASIC for the model (or group of models); and at that point you don't need Nvidia. I suppose this scenario will happen in various moments at different levels.
- hylaride 1mo agoIt's also hard to predict how much money will be burned going down wrong avenues. The internet was the future, but it took a lot of failed companies to eventually land on a sustainable model that brought us the giants we have today. Railways were also the future, but that didn't stop a rush to build out (often subsidized) lines that were ultimately uneconomical (either because they were corrupt or the planned settlements never arrived). If AI is similar, then there's going to be a long slowdown on compute spend until the surplus is worked through. A good historical analogy could be the fiber optic buildouts of the late 1990s. The demand for data never really went down much, but the industry eventually commodified and took down some large companies (Nortel, especially)
- galaxyLogic 1mo agoI rhink there's a difference with AI because -- it brings true value because you pay for the tokens, you only pay for what you use. That is true value. Compare to just paying for an internet connection, you have bandwidth but not sure what you can do with it that is valuable. Let's say you use AI to produce software. There;s no limit as to how high the quality you want your software to have. And how fast you want your project to be complete. There's plenty of room for higher quality, and more performant AI. As AI becomes chepaer people will use more of it, they're not going to say "We have enough AI". Compare to railroads. Yes you pay for the distance travelled but there's a limit to how much people wwill want to travel, how it will benefit them.
- aurareturn 1mo agoIt is entirely possible that 5 years from now, there's >100x LLM inference going on - but demand for AI chips (including memory) is only 2x or less. I doubt it. If LLM inference efficiency is 50x better than today, then there could be 1000x increase in inference volume due and we'll end up needing even more chips. Jevons paradox should win out for a long time for AI. When internet connections got faster than 56k modems, we didn't use the same amount of bandwidth but faster. We used more bandwidth doing things like 4k streaming. I see the same in AI inference. If AI inference is that much more efficient, it will just enable more use cases for AI. See for example, internet traffic over time: https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcS779VSUSjboPrTssAmLebYyENSW1c5VlfozbP8immqgMSIsncUyQStyf3o&s=10 https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcS779VS... Even after so many years, internet traffic continues to grow at an increasing rate.
- maerF0x0 1mo agoPlus on top of that 1st and 2nd order can be correct, but then the price is too high, meaning people lose money even if correct about the future, but over pay for it.
- FuriouslyAdrift 1mo agoThey also have to be feeling the heat of the ASIC vendors. AMD just acquired Taalas and they work with Cerebras all the time on special projects. ASICs outgun nVidia's chips by an order of magnitude.
- wmf 1mo agoNot really. The GPU+LPU combination is going to be pretty good once it comes out.
- whatever1 1mo agoEach of the hyperscalers has put like 250B each in the last year for infra. That means that they need to be writing AI profits to the tune of 20B per year just to keep up with the cost of the cash they burned. We are not there. But they better figure it out soon. The cash flows dried up, and everyone is taking debt to support the capex. Google for the first time in its public history is cash flow negative. Amazon too.