8 ms·
I think efficiency is unlikely to result in lower demand for compute, instead more useful compute per watt increases the value of that compute; and we are not g
by mattnewton 1mo ago
I think efficiency is unlikely to result in lower demand for compute, instead more useful compute per watt increases the value of that compute; and we are not going to run out of economically useful things to do with it anytime soon on the demand side.
The harder thing to forecast for me is if we hit a wall on increasing efficiency, either on the model weights side or silicon side, with current approaches. If we have to switch to something like burning the model weights into silicon to continue to make gains, then the current math on general purpose accelerators might be upside down.
- ekunazanu 1mo agoI agree; I don't think there's any reason to assume Jevons paradox won't apply. > If we have to switch to something like burning the model weights into silicon to continue to make gains I think that's already being considered semi-seriously [0][1] [0] https://taalas.com/products/ https://taalas.com/products/ [1] https://ir.amd.com/news-events/press-releases/detail/1296/amd-acquires-taalas-to-advance-compute-solutions-for-rapidly-growing-ai-inference-market https://ir.amd.com/news-events/press-releases/detail/1296/am...
- kurthr 1mo agoWhat's really interesting is that if you scale it to higher densities (eg 3nm and stacked die) with ComputeInMemory for fp8 you can reasonably start to fit 30B-70B models. With MoE and multiple stacked die, just like HBM, you could fit an open weight near frontier 1T model (like GLM5.2) at similarly much lower power <10kW and high token rates >2ktps. For running a bunch of agents where fill rate and speed/latency are important it may not matter that you're 6-18mo behind on weights. The process for the chip design could be largely automated, and new silicon pumped out as new weights are available (with a 3-6mo delay).
- adrianN 1mo agoWhen efficiency reaches the point where local models on consumer hardware are good enough, demand for cloud tokens could rapidly shrink.
- pseudosavant 1mo agoIt could be quite a while before we reach that point though. 5+ years easily. I've been keenly interested in the ability to run local models, but the hardware is just not there. Consumer RAM speeds and capacity will have to significantly increase before local models will be able to perform as well as even the lowest end GPT-5.6 Luna model. This is on the backdrop of RAM becoming prohibitively expensive. And without the speed and quantity of RAM, it becomes impossible to generate tokens at interactive speeds, regardless of model. There is a fundamental dependency between calculating all of the active params with the given RAM speed. Even with a model that has been quantized all the way down to Q4, the DGX/RTX Spark chip with 128GB of RAM can only generate ~18 tokens/sec for a MoE model with only 30B active parameters. There haven't been any broadly useful models below 30B active parameters. And that is for a $5000+ piece of hardware that will be one of the best for running on-device models. I really want to buy instead of rent my AI, but the economics are truly terrible.
- bryanlarsen 1mo agoVery few consumers are going to spend multiple thousands of dollars to save $10 per month. Companies absolutely will to save hundreds per month per employee, but that's not consumer hardware.
- HDBaseT 1mo agoMany gamers already spend $1000+ on a GPU. If you can integrate AI accelerators into consumer cards (you can), you can have local AI for "reasonably" cheap. This is Nvidia's long term goal if you listen to what Jensen has to say. The limitation is entirely on memory right now. Just a few years ago we could of been strapping 80-100GB to cards for under $200 (BoM).
- adrianN 1mo agoWell if the trend that the comment further up in this thread claimed continues and compute requirements keep dropping exponentially then perhaps in a few years you can have today’s frontier performance on the normal laptop you already have on your desk anyway.
- m463 1mo ago> think efficiency is unlikely to result in lower demand for compute You can't save yourself rich.
- ryukoposting 1mo ago> efficiency is unlikely to result in lower demand for compute, instead more useful compute per watt increases the value of that compute I buy that. Jevon's Paradox, sure. > and we are not going to run out of economically useful things to do with it anytime soon on the demand side This I don't buy. Not fully, at least. Whether or not there's demand for LLMs in some particular field is one thing, whether or not there is a sustainable business model to be built out of that demand is another thing entirely. There is a staggering amount of money pouring into startups looking for novel use cases for LLM-based agents. As usual, 99% of them will fail, but those other 1% are going to have to look harder and harder to find a novel use case that can actually be served profitably. First of all, there's only so many places where a chatbot is going to sell. But, that also seems to be the only interface anyone can come up with that allows a user to steer an agent through a long-running task reliably. I'd love to be proven wrong here. Also, if current trends plateau and large datacenters are still needed for complex tasks, that would stimy growth of LLM usage across entire industries. But, if present trends continue, then local inference will become feasible for most tasks. That would lower the barrier to entry across tons of heavily-regulated and/or cost-sensitive industries. But, widespread local inference will almost certainly come with a painful market correction centered around hyperscalers, which would itself dry up the pool for ventures into new markets.
- the_sleaze_ 1mo agoReplace "chat-bot" with "Human Being" because the models I've been using are significantly better than 90% of the human-chat-bots that I must talk to on the phone while scheduling and coordinating my internet installation for example. Now for every human replacement, that is 1 unit less of communication and bureaucratic burden (HR, middle management etc) that the org requires.
- ryukoposting 1mo ago> the models I've been using are significantly better than 90% of the human-chat-bots that I must talk to on the phone while scheduling and coordinating my internet installation for example. I guess I don't know what to say except that my experience is the polar opposite of yours. I moved to a new state at the beginning of the year. Needed a new doctor, needed to schedule apartment tours, needed to talk to my employer about insurance and relocation stuff, etc etc. Lots of chatbots, a handful of humans. Humans consistently did what I needed them to do, the chatbots just didn't. I could list examples but I'd be typing all night. And yknow what, my one call with Comcast to get my internet set up was downright pleasant. The rep was knowledgeable and a good conversationalist.