7 ms·
I think it is all well and good, but the most affordable option is probably still to buy a used MacBook with 16/32 or 64 GB (depending on the budget) unified me
by rthnbgrredf 1y ago
I think it is all well and good, but the most affordable option is probably still to buy a used MacBook with 16/32 or 64 GB (depending on the budget) unified memory and install Asahi Linux for tinkering.
Graphics cards with decent amount of memory are still massively overpriced (even used), big, noisy and draw a lot of energy.
- jibbers 1y agoGet an Apple Silicon MacBook with a broken screen and it’s an even better deal.
- ivape 1y agoIt just came to my attention that the 2021 M1 Max 64gb is less than $1500 used. That’s 64gb of unified memory at regular laptop prices, so I think people will be well equipped with AI laptops rather soon. Apple really is #2 and probably could be #1 in AI consumer hardware.
- jeroenhd 1y agoApple is leagues ahead of Microsoft with the whole AI PC thing and so far it has yet to mean anything. I don't think consumers care at all about running AI, let alone running AI locally. I'd try the whole AI thing on my work Macbook but Apple's built-in AI stuff isn't available in my language, so perhaps that's also why I haven't heard anybody mention it.
- ivape 1y agoPeople don’t know what they want yet, you have to show it to them. Getting the hardware out is part of it, but you are right, we’re missing the killer apps at the moment. The very need for privacy with AI will make personal hardware important no matter what.
- mycall 1y agoTwo main factors are holding back the "killer app" for AI. Fix hallucinations and make agents more deterministic. Once these are in place, people will love AI when it can make them money somehow.
- croes 1y agoYou can’t fix the hallucinations
- herval 1y agoHow does one “fix hallucinations” on an LLM? Isn’t hallucinating pretty much all it does?
- kasey_junk 1y agoCoding agents have shown how. You filter the output against something that can tell the llm when it’s hallucinating. The hard part is identifying those filter functions outside of the code domain.
- dotancohen 1y agoIt's called a RAG, and it's getting very well developed for some niche use cases such as legal, medical, etc. I've been personally working on one for mental health, and please don't let anybody tell you that they're using an LLM as a mental health counselor. I've been working on it for a year and a half, and if we get it to production ready in the next year and a half I will be surprised. In keeping up with the field, I don't think anybody else is any closer than we are.
- tptacek 1y agoWait, can you say more about how RAG solves this problem? What Kasey is referring to is things like compiling statically-typed code: there's a ground truth an agent is connected to there --- it can at least confidently assert "this code actually compiles" (and thus can't be using an entirely-hallucinated API. I don't see how RAG accomplishes something similar, but I don't think much about RAG.
- dotancohen 1y ago> People don’t know what they want yet, you have to show it to them Henry Ford famously quipped that had he asked his customers what they wanted, they would have wanted a faster horse.
- airtonix 1y ago[dead]
- estimator7292 1y agoWe've shown people so many times and so forcefully that they're now actively complaining about it. It's a meme. The problem isn't getting your Killer A I App in front of eyeballs. The problem is showing something useful or necessary or wanted. AI has not yet offered the common person anything they want or need! The people have seen what you want to show them, they've been forced to try it, over and over. There is nobody who interacts with the internet who has not been forced to use AI tools. And yet still nobody wants it. Do you think that they'll love AI more if we force them to use it more?
- ivape 1y agoAnd yet still nobody wants it. Nobody wants the one-millionth meeting transcription app and the one-millionth coding agent constantly, sure. It a developer creativity issue. I personally believe the creativity is so egregious, that if anyone were to release a killer app, the entirety of the lackluster dev community will copy it into eternity to the point where you’ll think that that’s all AI can do. This is not a great way to start off the morning, but gosh darn it, I really hate that this profession attracted so many people that just want to make a buck. ——- You know what was the killer app for the Wii? Wii Sports. It sold a lot of Wiis. You have to be creative with this AI stuff, it’s a requirement.
- deleted 1y ago[deleted]
- wkat4242 1y agoM1 doesn't exactly have stellar memory bandwidth for this day and age though
- Aurornis 1y agoM1 Max with 64GB has 400GB/s memory bandwidth. You have to get into the highest 16-core M4 Max configurations to begin pulling away from that number.
- wkat4242 1y agoOh sorry I thought it was only about 100. I'd read that before but I must have remembered incorrectly. 400 is indeed very serviceable.
- deleted 1y ago[deleted]
- benreesman 1y agoRyzen AI 9 395+ with 64MB of LPDDR5 is 1500 new in a ton of factors and 2k with 128. If I have 1500 for a unified memory inference machine I'm probably not getting a Mac. It's not a bad choice per se, llama.cpp supports that harware extremely well, but a modern Ryzen APU at the same price is more of what I want for that use case, with the M1 Mac youre paying for a Retina display and a bunch of stuff unrelated to inference.
- anonym29 1y agoNot just LPDDR5, but LPDDR5X-8000 on a 256-bit bus. The 40 CU of RDNA 3.5 is nice, but it's less raw compute than e.g. a desktop 4060 Ti dGPU. The memory is fast, 200+ GB/s real-world read and write (the AIDA64 thread about limited read speeds is misleading, this is what the CPU is able to see, the way the memory controller is configured, but GPU tooling reveals full 200+ GB/s read and write). Though you can only allocate 96 GB to the iGPU on Windows or 110 GB on Linux. The ROCm and Vulkan stacks are okay, but they're definitely not fully optimized yet. Strix Halo's biggest weakness compared to Mac setups is memory bandwidth. M4 Max gets something like 500+ GB/s, and M3 Ultra gets something like 800 GB/s, if memory serves correctly. I just ordered a 128 GB Strix Halo system, and while I'm thrilled about it, but in fariness, for people who don't have an adamant insistence against proprietary kernels, refurbished Apple silicon does offer a compelling alternative with superior performance options. AFAIK there's nothing like Apple Care for any of the Strix Halo systems either.
- jtbaker 1y agoThe 128 GB Strix Halo system was tempting me, but I think I'm going to hold out for the Medusa Point memory bandwidth gains to expand my cluster setup. I have a Mac Mini M4 Pro 64GB that does quite well with inference on the Qwen3 models, but is hell on networking with my home K3s cluster, which going deeper on is half the fun of this stuff for me.
- ivape 1y agoIt’s not better than the Macs yet. There’s no half assing this AI stuff, AMD is behind even the 4 year old MacBooks. NVDIA is so greedy that doling out $500 dollars will only you get you 16gb of vram at half the speed of a M1 Max. You can get a lot more speed with more expensive NVDIA GPUs, but you won’t get anything close to a decent amount of vram for less than 700-1500 dollars (well, truly, you will not get close to 32gb even). Makes me wonder just how much secret effort is being put in by MAG7 to strip NVDIDA of this pricing power because they are absolutely price gouging.
- seanmcdirmid 1y agoI recently got an M3 Max with 64g (the higher spec max) and ts been a lot of fun playing with local models. It cost around $3k though even refurbished.
- giancarlostoro 1y agoYou dont even need Asahi, you can run comfy on it but I recommend the Draw Things app, it just works and holds your hand a LOT. I am able to run a few models locally, the underlying app is open source.
- mrbonner 1y agoI used Draw Thing after fighting with comfyui.
- croes 1y agoWhat about AMD Ryzen AI Max+ 395 Mini PCs with upto 128GB unified memory?
- evilduck 1y agoTheir memory bandwidth is the problem. 256 GB/s is really, really slow for LLMs. Seems like at the consumer hardware level you just have to pick your poison or what one factor you care about most. Macs with a Max or Ultra chip can have good memory bandwidth but low compute, but also ultra low power consumption. Discrete GPUs have great compute and bandwidth but low to middling VRAM, and high costs and power consumption. The unified memory PCs like the Ryzen AI Max and the Nvidia DGX deliver middling compute, higher VRAMs, and terrible memory bandwidth.
- codedokode 1y agoBut for matrix multiplication, isn't compute more important, as there are N³ multiplications but just N² numbers in a matrix? Also I don't think power consumption is important for AI. Typically you do AI at home or in the office where there is lot of electricity.
- evilduck 1y ago>But for matrix multiplication, isn't compute more important, as there are N³ multiplications but just N² numbers in a matrix? Being able to quickly calculate a dumb or unreliable result because you're VRAM starved is not very useful for most scenarios. To run capable models you need VRAM, so high VRAM and lower compute is usually more useful than the inverse (a lot of both is even better, but you need a lot of money and power for that). Even in this post with four RPis, the Qwen3 30 A3B is still an MOE model and not a dense model. It runs fast with only 3B active parameters and can be parallelized across computers but it's much less capable than a dense 30B model running on a single GPU. > Also I don't think power consumption is important for AI. Typically you do AI at home or in the office where there is lot of electricity. Depends on what scale you're discussing. If you want to get similar VRAM as a 512GB Mac Studio Ultra with a bunch of Nvidia GPUs like RTX 3090 cards you're not going to be able to run that on a typical American 15 AMP circuits, you'll trip a breaker half way there.
- Aurornis 1y ago> and install Asahi Linux for tinkering. I would recommend sticking to macOS if compatibility and performance are the goal. Asahi is an amazing accomplishment, but running native optimized macOS software including MLX acceleration is the way to go unless you’re dead-set on using Linux and willing to deal with the tradeoffs.
- benreesman 1y agoIf Moore's Law is Ending leaks are to be believed, there are going to be 24GB GDDR7 5080 Super and maybe even 5070 Super Ti variants in the 1k (MSRP) range and one assumes fast Blackwell NVFP4 Tensor Cores. Depends on what you're doing, but at FP4 that goes pretty far.
- nullsmack 1y agoThe mini pcs based on AMD Ryzen AI Max+ 395 (Strix Halo) are probably pretty competitive with those. Depending on which one you buy it's $1700-2000 for one with 128GB RAM that is shared with the integrated Radeon 8060S graphics. There's videos on youtube talking about using this with the bigger LLM models.