6 ms·
Apparently GLM 5.2 is 753B parameters [1], what kind of hardware are people using to run this locally? [1] https://huggingface.co/zai-org/GLM-5.2 https://huggi
by bArray 3mo ago
Apparently GLM 5.2 is 753B parameters [1], what kind of hardware are people using to run this locally?
[1] https://huggingface.co/zai-org/GLM-5.2 https://huggingface.co/zai-org/GLM-5.2
- crocowhile 3mo agofollow antirez - https://x.com/antirez/status/2071173841175363905?s=20 https://x.com/antirez/status/2071173841175363905?s=20
- JamesSwift 3mo agoThats quantized
- nozzlegear 3mo agohttps://xcancel.com/antirez/status/2071173841175363905 https://xcancel.com/antirez/status/2071173841175363905
- anentropic 3mo agoIt's a nice technical achievement but looks unusably slow for actual work
- dakolli 3mo ago8 X RTX6000. It will run you around 80-100k to get started with a model at this size with decent tps.. Don't worry though, open source evangelists will tell you that these will be running on your phone in the next 3 years. For $100k you could run this model 24/7 through open router with 10 concurrent sessions at 50tps for a decade and have money left over for a vacation. There's no point in investing this type of money in local models unless you have a business where you're already paying for many employee's individual token usage.
- 8note 3mo agoyou can however, have fun with it. oil workers buy 100k trucks they do not-much with. why not a 100k in computer?
- dakolli 3mo agoSure, If you want to light money on fire for entertainment, more power to you. There's probably worse ways to light 100k on fire. If I have an extra 100k laying around it's going to my family though.
- Ken_At_EM 3mo agoI can't help but ask where this comment came from, you must have some exposure..
- CamperBob2 3mo agoIt is so easy to spend $100K on a pickup truck these days, it's not even funny.
- SV_BubbleTime 3mo agoFactory F350 Platinum is at least 90k sticker.
- hedora 3mo agoYet Ford claims it is impossible to sell any pickups for > $60K, so they killed the lightning. I assume (since they claim they are selling the batteries to AI data centers), they’ll produce some sort of EV >= F150 once the bubble pops, and we get a new president.
- SV_BubbleTime 3mo agoAutomotive EE here… every other decision about vehicles is about emissions. CAFE, the reason that a company releases X model is that they can then sell more Y models that get worse mileage. EV is a separate thing. Vastly overmarketed for the technology as it exists today.
- rekttrader 3mo agoOr you have data that HIPAA, GDPR, PII, or have to care about the concern of others training on your data.
- dakolli 3mo agoThat too.
- krackers 3mo agoWould you be better off pooling that money with some hackerspace group and then setting up shared inference infra, so that way you at least get better utilization?
- KaoruAoiShiho 3mo agoAnd before you know it, you invented some openrouter provider from first principles...
- janalsncm 3mo agoRight. For example you will need to figure out how to share it and who maintains it.
- aetch 3mo agoYou can then rent spare capacity out to people on a subscription or token basis ….wait
- dist-epoch 3mo ago> 50tps for a decade assuming demand doesn't keep on increasing. even google has trouble having enough capacity apparently.
- Aurornis 3mo ago> 8 X RTX6000. It will run you around 80-100k to get started 8 x RTX6000 GPUs cost $100,000 alone. You then need to build a system that can support those GPUs with enough PCIe lanes through a PCIe switch. It's going to be $120K to $150K to build or buy a system to run this.
- CamperBob2 3mo agoYou can run the NV4FP quant with 8x RTX6000 cards at 50-75 tps output, but not (practically speaking) the OEM FP8 version. You will learn more about PCIe than you ever wanted to know. The real gangstas are running 16x RTX6000s. Too rich for my blood, and the NV4FP quant doesn't seem to be that much worse.
- InvertedRhodium 3mo agoDepends how much you value privacy and running uncensored models. Personally, I’m waiting for hardware to hit the secondary market before I buy something to run unquantized models like GLM. But I have no doubt that I will, at some point.
- wonnage 3mo agoYeah, the neoclouds and hyperscalers are taking massive losses right now, self hosting is basically signing yourself up to do the same. There are philosophical reasons to do so but it’s a terrible economic decision
- KetoManx64 3mo agoAs an individual I do not need the whole model. I don't need the model to have knowledge of the rain history of Algeria nor how many colors are in the Russian flag. Once they start trimming down the excess and making them field focused they will run just fine on people's individual devices.
- JumpCrisscross 3mo ago> I do not need the whole model. I don't need the model to have knowledge of the rain history of Algeria nor how many colors are in the Russian flag Isn’t the performance gap between quantized and full models indicative that even if you aren’t using it directly, the model knowing the colors in the Russian flag does have something to do with the intelligence you demand?
- KetoManx64 3mo agoDo quantized models specifically prune out specific knowledge? I think they just compress things down but they're still in there. You'd most likely need to do that when you're doing the initial model training, but I'm not expert.
- JumpCrisscross 3mo ago> they just compress things down but they're still in there The compression is almost certainly in part specific knowledge getting fuzzed.
- DennisP 3mo agoYeah, but it's everything getting fuzzed, including the parts you care about.
- JumpCrisscross 3mo agoSure. There is a legitimate question around whether one can selectively excise “useless” knowledge. My guess is you can’t. The act of learning it encodes both the act of learning and the knowledge per se. The former is the power of the LLM. (I personally force mine to double check everything instead of going off memory.)
- Ldorigo 3mo agoHow do the economics of your statement work out? Clearly inference providers don't have a time to ROI of 10 years on their hardware costs; and that's without even taking ongoing energy costs into account. What's missing here?
- ac29 3mo agoThe inference providers are running batch sizes much larger than 10
- dakolli 3mo agohttps://aimultiple.com/gpu-benchmark https://aimultiple.com/gpu-benchmark concurrency
- kingstnap 3mo agoOutput tokens are actually kinda expensive for the provider. The input cache hit tokens are incredibly cheap for them, (incredibly high margin too, except for deepseek). And input tokens are in the middle. Input tokens can be processed very efficiently. Also his math is wrong. $100k gets you 22.7B output tokens at $4.4/M which is how much GLM 5.2 costs. At 500/s 22.7B is just 500 days. Or about 1.54 years. Which is much less then the life of the hardware.
- bandrami 3mo agoInference providers have been getting a firehose of investor cash to keep the chips running (and are looking around very nervously as that firehose starts to sputter).
- AussieWog93 3mo ago>Don't worry though, open source evangelists will tell you that these will be running on your phone in the next 3 years. Not sure if you're being sarcastic, but I can run a quantised version of Gemma or Qwen on my 16GB M1 Macbook Pro that beats GPT-4 from 2023 hands-down. I wouldn't be surprised if, in another 3 years, you'd be able to run something as powerful as Opus 4.5 or GLM-5.2 on standard consumer hardware - say a 32GB/64GB M7 Pro. I also wouldn't be surprised if, 3 years after that, cheaper hardware and improved model efficiency means that there's a much smaller gap between what you can run on a consumer CPU (which, with memory prices coming down, could look like a 256GB M9 or M10 Pro) and $100k GPU cluster.
- marcus_holmes 3mo agoThis is clearly where the industry is going, imho. Everyone who is playing with LLMs wants a laptop with enough grunt to run a decent model locally. We've been sat with basically the same PC specs for ~20 years - our current specs are within an order of magnitude of the ones we could buy back in 2010. This is not really constrained by tech, as we could have much, much, larger machines. It's more because there's no mass demand for much, much, larger machines - if it's big enough to run Office apps or VSCode then you're good to go. The exponential growth we saw in the 90's was driven as much by software demand as it was by hardware development. I can see the next 10 years produce the same kind of push for larger machines that the 90's did. And we should probably expect the same kind of standards churn as our existing technologies for storage, memory, etc, don't scale up enough and new technologies become worth developing because there's demand for them.
- byzantinegene 3mo agomy only concern if the same specs today would cost 10x more given the trajectory of the growth of memory prices lately.
- marcus_holmes 3mo agoI think this is where the new technology comes in. There is demand for 10x (or 1000x) the memory that we're using at the moment, so someone/something will satisfy that demand. We haven't had that demand up until now, because 16Gb was a perfectly reasonable amount of memory that could run pretty much anything, and if that won't then 32Gb will. There was zero demand for 16Tb memory machines because no-one had any application for that much memory. Now that's changing, and there is demand for that much, so we'd expect to see that being made available. But the existing tech we're using for 16Gb probably isn't going to scale to 16Tb at a reasonable price point. And the price point is relatively inelastic - people are used to paying <$5K for their computers, and they're not going to go much above that. You'll get early adopters paying $10K or more for a machine that large, but not the early majority. And even then, obviously, $10K is not going to buy you a 16Tb memory machine. So there's room for a new technology to come in, where there wasn't previously. This is what happened all through the 90's, and we churned through a bunch of standards and technologies to try and keep up with demand.
- DrScientist 3mo agoGiven GLM is open weight - all you need is one company to take the taalas approach ( model on hardware ), and you're sorted right? https://taalas.com/products/ https://taalas.com/products/
- akie 3mo agoYeah I completely agree. But this is much larger model than the 8B one they put on a chip, so that's probably an engineering challenge for now. Also, how expensive would it be?
- DrScientist 3mo agoNo idea - AI tells me under 30 dollars per unit for the ROM with development costs in the low 10's of millions. If that's anywhere near right then it seems like a no brainer.
- Rekindle8090 3mo ago[dead]
- deleted 3mo ago[deleted]
- kccqzy 3mo agoRun quantized versions. https://unsloth.ai/docs/models/glm-5.2 https://unsloth.ai/docs/models/glm-5.2
- Retro_Dev 3mo agoI ran it on my laptop, which is a Lenovo Legion 5i (think 32 GB RAM, 4060 w/ 8 GB VRAM, you get the picture). It was a quantized model (otherwise it would not fit on my NVMe 1TB drive) at 4 bits per weight - UD_Q4_K_XL. It ran at about 12 seconds per token (not tokens per second). A fun project, but not worth it. I used 4096 tokens of context cache, and I ran it with llama.cpp - as it supports memory mapping. Because the whole thing could obviously not fit in RAM, I was curious how much it would need to stream from SSD. The answer? For a simple 4 sentence description of who it was, about 1.5 TiB was streamed from disk.
- bArray 3mo agoThank you for sharing. 1.5TB of streamed data at 12 seconds per token on a high end consumer laptop is a pretty high requirement - I can only imagine how much that cost to train. I don't know how running this model could be cost effective for anybody.
- Retro_Dev 3mo agoIndeed - definitely not cost effective to run it on this laptop LOL. It makes me wonder how fast we could run the model if we could fit the weights entirely within CPU cache (assuming a whole ton of CPUs with low latency & high speed IO of course).
- scosman 3mo agoshort answer: they mostly aren't A few people are running highly quantized models with limited context windows. It's still impressive, but not the benchmark level intelligence. Very few people could afford a rig for reasonable local performance at a reasonable quant, at full context size. The antirez example is 2.6bit quant, 32k context, and few tokens per second... on a ~$7000 MacBook M5 (new RAM pricing).
- deleted 3mo ago[deleted]