6 ms·
Unified Memory, Explained: Why Mini PCs Can Run 70B Models a Big GPU Can't
- _davide_ 2mo agoIf compute is not the bottleneck, memory is easy-ish to produce (the hard part is mostly on the fab side); what stops a Chinese NVIDIA (huawei) from being 10x cheaper?
- WhyNotHugo 2mo agoI think it's mostly the ramp-up time, but ChangXin Memory Technologies (CXMT) is basically aspiring to do just this.
- selectodude 2mo agoMaking memory is easy. Packaging that memory within a few millimeters of a piece of silicon using TSVs and maintaining signal integrity on a 1024 bit bus is really, really hard. LLMs aren’t all that compute constrained or even memory constrained. It’s just that pushing dozens of terabits per second through a piece of silicon is a physics problem.
- erkt 2mo agoUhh the 5090 alone is double the cost of their quoted PC prices.
- LoganDark 2mo agoJust get a used one. I picked up a used 3090 a couple years back for $800. Smelled like tobacco but works just fine.
- LoganDark 2mo agoCan't really run it as well, though. My "mini PC" is an M4 Max with 128GB of unified memory and the memory bandwidth is still sorely lacking for inference (although it's far better than any non-unified consumer architecture!).
- bahmboo 2mo agoTo be fair it's "only" half the throughput of a 4090 and a third of an RTX 6000. Significant but not an order of magnitude.
- lowbloodsugar 2mo agoAn old ada Rtx 6000 maybe. A Blackwell RTX Pro 6000 is an order of magnitude faster and has 96gb.
- bahmboo 2mo agoThat's not what I'm seeing. It is much faster but not an order of magnitude. Not trying to be pedantic, only setting expectations. "The Blackwell RTX PRO 6000 provides up to 1,792 GB/s of memory bandwidth, while the 40-core Apple M5 Max tops out at 614 GB/s"
- lowbloodsugar 2mo agoSorry, thought we were talking about tokens. M5 Max is great for bandwidth and I’m looking forward to seeing what Apple does for AI inference in the M7. The 6000 kills everything else when it comes to TTFT and tokens/s.
- bahmboo 2mo agoFor sure. Clearly Nvidia mops the floor with the competition. I'm looking forward to M6/M7 and to see if Apple wants a bigger piece of the pie.
- entrope 2mo agoThose are the ratios for memory bandwidth, but the GPUs have a much higher ratio for compute, and that affects prefill rate / TTFT, right?
- LoganDark 2mo agoFor local inference, the difference between 25t/s and 70t/s is a lot. For some models I struggle to even reach 15t/s. And "some models" aren't even large models, Gemma 4 13b has this issue for some reason. For stuff like Qwen3.6-27B I can hardly reach 10t/s, even with fully custom inference made by Fable 5!
- OutOfHere 2mo agoLet's also ensure the SSD doesn't age prematurely.
- craftkiller 2mo agoI was under the impression that when you're streaming the weights from disk because the full model won't fit in memory, that it is solely reading from the SSD, not writing, so it wouldn't be causing wear on your SSD.
- 1000100_1000101 2mo agoYou'd need your OS to support, and be configured to use, a disk mounting option that disables file access timestamps, otherwise reads ARE writes.
- trinix912 2mo ago...and disable swap.
- wtallis 2mo agoDo any operating systems update file access time after every read operation instead of just at fopen?
- 1000100_1000101 2mo agoBased on the ~10x performance gain from disabling access timestamps in VxWorks' HRFS, I'd guess it does. I have no idea if that's common or rare.
- giantrobot 2mo agoIt is and it doesn't. You only get into disk writes if the system starts paging out to disk.
- OutOfHere 2mo ago
- Havoc 2mo agoThink future generations of AMD could get quite interesting. They’re no doubt seeing people whining about mem throughput specifically
- throwa356262 2mo ago"Can't" is not really correct. Nowadays, specially with MoE models you can run parts of the model on GPU and still get some speed up.
- reinitctxoffset 2mo agoThis is a very understandable misconception that I wouldn't blame anyone for having but MoE is actually terrible for inference in most any local LLM / home lab scenario. MoE is popular because it's cheap to train, but because most modern routing needs the previous layer's activations (except at the very beginning) it winds up being just this side of impossible to pipeline / prefetch without all the experts resident. Plus the grouped GEMM kernels have terrible support on any card in most people's house, it's just really unwieldy. Dense models are very straightforward to share/pipeline because you know all the shapes and geometry up front, that's the inference friendly option. MoE sells a lot of HBMe3.
- throwa356262 2mo agoI thought it was easier to find a possible cut in a MoE where the amount of data transferred between layers is very small (kilo bytes) while in dense architecture this is much harder?
- ReptileMan 2mo ago[flagged]
- cheevly 2mo agoMaybe like… focus on the actual content instead of perceived writing patterns? Crazy I know.
- glimshe 2mo agoThis. I don't particularly like the LLM writing style, but we've read a ton of very poorly written texts over the year with no complaints. LLMs are average writers with an annoying style, but not bad writers. If the content is good, I don't care if a LLM wrote it.
- jazzyjackson 2mo agoIt’s a flag that no care went into it. Web surfing involves lots of little decisions following cues of “is this worth my time”
- toast0 2mo agoIt's worse than no care... I read an article recently that was written with no care and it was refreshing. Putting an LLM on it means you care to make it look nice, but not enough to actually do it. Why bother?
- deleted 2mo ago[deleted]
- lukeinator42 2mo agoThe writing style is this staccato LLM-like style that is difficult to read as it has zero flow and meaningless sentences. Like what does the second sentence even mean? Is it even a sentence? "The roofline math, the prompt-processing catch, the NPU red herring, and the owner-measured speeds."
- lowbloodsugar 2mo agoThe current “big GPU” has 96gb of memory, but that’s not a “consumer GPU” apparently, while a $5000 Spark is a “consumer PC” I guess. In any case you’re probably better off running a large open weights model on the cloud.
- amelius 2mo agoDo unified memory CPUs suffer from the same memory shortages as normal memory? I guess they're just welding the memory to the CPU chip, but still curious.
- _davide_ 2mo agoThey are usually the same family, LPDDR is used for amd and macs, but the fabs are the same as the most expesive HBM memory, if they have a choice they are going to produce the ones that they can sell for more $$.
- bahmboo 2mo agoYes. The memory is just located very close to the cpu with wires "welded" directly to it. This allows the memory to be run as fast as possible but it's still a RAM component. The cache parts of memory are on the CPU itself but they are on the order of MB not GB.
- wtallis 2mo ago> I guess they're just welding the memory to the CPU chip, but still curious. Unified memory is more of an architectural and performance characteristic, and does not imply much about the physical layout of the machine. Most unified memory PCs not from Apple don't have the memory on the same package as the SoC. For stuff like AMD Strix Halo and NVIDIA DGX Spark, it's just standard LPDDR packages soldered on the motherboard in the general vicinity of the SoC, and the only difference from mainstream laptops for the past decade+ is that the memory bus is twice as wide.
- _davide_ 2mo agoI'm writing my own inference engine for Strix Halo and the same model. I already have 30%+ performance plus a more graceful decay over long contexts; that said, their point stands: memory bandwidth is what you really want.
- bdcravens 2mo ago> Put two machines on a desk, each about $2,000. One is a tower with an NVIDIA RTX 5090: 32GB of the fastest consumer memory ever shipped, 1,792 GB/s. The other is a mini PC the size of a paperback, an AMD Ryzen AI Max+ 395 "Strix Halo" box with 128GB of soldered memory at roughly 256 GB/s. Doesn't change the conclusions of the article, but each of those machines is more like $4k+ https://www.microcenter.com/product/711961/amd-ryzen-ai-halo-developer-platform-linux-os https://www.microcenter.com/product/711961/amd-ryzen-ai-halo...
- cocodill 2mo agoFor some reason, this reminds me of my last shared memory system. It was an Athlon XP 1800+ with VIA ProSavage back around 2002. It was just barely able to run CS 1.6.
- geon 2mo agoReally? I broke my Geforce2 MX, so I had to make do without a graphics card for a couple of weeks. I think halflife ran ok in software mode on my Athlon XP 1700+. I might be misremembering though. Perhaps I scavenged some basic pci card, but that should still have been worse than the ProSavage.
- danbruc 2mo agoWhy would a RTX 5090 with 32 GB not be able to deal with a 40 GB model? Is there anything preventing me from swapping the weights that do not fit into VRAM in and out of RAM? PCIe 5.0 x16 should max out around 64 GB/s, so slower than the unified memory machine, but at least it should be possible.
- buckle8017 2mo agoIt's slower than the 4:1 ratio would imply, but it does indeed work. Things get really slow if the model doesn't for in vram + ram and you have to go from disk to ram to vram.
- searealist 2mo agoThere are two phases to LLMs: 1) prefill 2) decode For prefill, you are compute bound, and it is trivial to batch multiple input tokens together. When using cpu offload, software like llama.cpp will batch weight uploads with tokens that need those weights and perform work on the GPU. It works very well. With a large batch size and pcie5 you can get prefill speeds close to having all weights on the GPU. For decode, you are bandwidth bound, and it is difficult to batch multiple output tokens together. There is no benefit to sending your weights to the GPU because even if it internally has insane bandwidth, you are still bottlenecked by system RAM (and adding a pcie5 upload would bottleneck it further). This is the number people usually talk about when they say they are getting a certain tk/s.
- hn_c 2mo ago> For decode, (...) it is difficult to batch multiple output tokens together. I think it's the other way around? The GPU has to stream gigabytes of active layer weights to compute the next token, so having a batch of next-token predictions sitting there on the GPU goingh through the layers makes better use of the bandwidth. At least that's what I observed on a Strix Halo, batching 4 predictions yields like 2-3x the total tps.
- searealist 2mo agoMTP can give you small batches, but it is still WAY smaller than the batches you can get with prefill, which is limited only by the number of input tokens you have (but has diminishing returns on performance). But: 1) It still makes no sense to upload the weights to the GPU with MTP as you are still bottlenecked by the weight upload. 2) I'm not sure MTP helps much with MoE models.
- vkaku 2mo agoI'm going to say this that we're not even close to the limits of what actually needs to be accomplished so at some point, memory will start needing better tiering for inference some day ....
- NortySpock 2mo ago"integrated graphics processor, using system memory" had its name dragged through the mud for decades. So we had to rebadge it to "unified memory". Curious if we'll ever see some old integrated graphics processor "hacked" to manage to handle 128 GB of allocated system RAM and be able to serve diffusion-LLMs at a decent rate on "old" hardware...
- officeplant 2mo agoafaik you can do that now on DDR4 platform Mini PC's that can handle 48gb or 64gb DIMM's.
- tim-tday 2mo agoI have trouble converting this article into actionable information.
- enbarca 2mo agoI got myself a NVIDIA DGX Spark to build on unified-memory hardware at home. Now running a health-AI node for all my longitudinal health data - insights are generated on the box so that raw data never leaks.