6 ms·
I was looking at $10k Mac Studio with M5 Ultra and 256 GB for local experiments, but then struggled to find what really good modern model I can fit into it. Yes
by ololobus 16d ago
I was looking at $10k Mac Studio with M5 Ultra and 256 GB for local experiments, but then struggled to find what really good modern model I can fit into it. Yes, it can run a good dense 27B at Q8 with plenty of context, but what beyond that? IIUC, some Deepseek flash variants at Q4 are also feasible, but I am not sure if the quality will be good. They also don’t run that fast, like about 30 t/s
So if I stay within 35B, especially MOE, my M5 Pro 64GB MBP can also run them well, and it can do plenty of other stuff too including gaming. While 256 GB with such RAM bandwidth and powerful GPU sounds like fun on paper, it doesn’t seem to be the next level compared to 64 GB
Really curious what people run on 256 GB Macs
- vadansky 16d agoSounds like it's worth waiting for M7 anyways, no point investing too much right now https://news.ycombinator.com/item?id=48676795 https://news.ycombinator.com/item?id=48676795
- Danox 16d agoAt least wait until someone gets one in hand and post a review, and if it does work decently well there probably is going to be a long backlog.
- robflynn 15d agoIs Apple still doing the extended return window for items receiving around the holiday? I know sometimes they allow items received near the end of November to be returned in January, so in theory if you time your early order right you can use it, test it, and even have some time to play with it before you return it.
- vablings 16d agoI feel like for localAI t/s is less of an issue. Just make a PRD and run a ralph loop. For big slogging projects like reverse engineering, or converting a codebase to a new language it actually doesn't matter if it takes a day or seven days.
- SMEbooop 16d agoYeah this is my experience. My 24GB 3090 + 64GB RAM takes a couple hours to crank out some code with largest Gemma 4 and Qwen3.8 models it can run But in the meantime I get dishes done, vacuum, flip laundry... etc etc Frontier models also seem in such a rush to emit anything they produce a mess that needs steering all day anyway While I have not tested it, it feels like my local setup going slower is better at producing code that works the first time as its not trying to look fast for marketing sake