7 ms·
Weights on HF here: https://huggingface.co/zai-org/GLM-5.3-Flash https://huggingface.co/zai-org/GLM-5.3-Flash I decided to take the plunge and get myself four
by mmastrac 21d ago
Weights on HF here: https://huggingface.co/zai-org/GLM-5.3-Flash https://huggingface.co/zai-org/GLM-5.3-Flash
I decided to take the plunge and get myself four sparks at a decent price (and bought the QSFP cables from AliExpress because they are literally 1/2 the price of Amazon), even knowing Apple was going to release new hardware and there's probably a spark 2 on the horizon. It looks like this is going to be a decent fit for what I need. I've been experimenting with a two-node DS4 and it's _good_ at some tasks, but it really just spins its wheels when it hits the limit of what it can reason through.
I can offload mundane/basic tasks to DS4 on two sparks, but I've been pushing it harder on some novel work and it just can't run on its own at all beyond a certain complexity level.
I would love to see an Opus-4.8-level local model but TBH I just haven't got there yet. The models I've tried so far _are_ good but they aren't able to solve tough technical challenges, regardless of harness/prompting/etc.
- kilroy123 21d ago> get myself four sparks at a decent price Wow, if you don't mind me asking. How and where?
- mmastrac 21d agoI bought 4x Asus GX10 with the 1TB option. I don't understand why, but it's the only model in the whole lineup that isn't priced insanely. They were briefly on sale with a $200-off coupon, but they show up on warehouse deals from time-to-time as well.
- swiftcoder 21d ago> it's the only model in the whole lineup that isn't priced insanely $4,000 isn't priced insanely? ye gads
- nijave 21d ago"For new hardware in 2026 with 128Gi of high-speed memory" Checked a couple days ago and looks like we're at about 3.5x 2020 memory prices (looking at just $/GB).
- a3w 21d agoI thought 4000 in sum. No wait, 4000 per, plus tax. Or EUR pricing to similar accord. Ouch.
- swiftcoder 21d agoYeah, that little cluster costs about the same as a brand-new Dacia Sandero.
- Bluestein 21d agoYeah, yeah. BUT, will the Sandero be ... load-bearing? :)
- KptMarchewa 21d agoIt can bear the load of a few people, at least.
- Bluestein 21d agoHey, I can at least say you will get more value (or, at least more predictable value) from a Dacia than from Anthropic's tokens: "Upgrade to [SUPER DUPER] for 5x the [TOKEN-SERF PACKAGE] token use!" Actual net work doable with/intelligence supplied by the [TOKEN-SERF PACKAGE]: Unknown. Fluctuating.-
- swiftcoder 21d ago> you will get more value (or, at least more predictable value) from a Dacia I picked up an older Dacia Sandero for cheap a few years back - it's the best money I've ever spent on a car, hands down. That car does not quit.
- PcChip 21d agoto save others from having to look up what that is like I did, it's a car
- esafak 21d agoYes, but it was $200 off!
- swatcoder 21d agoCompare to the cost of professional-grade tools in other trades and craft hobbies. Sure, $4000 can be a lot of if you're a casual hobbyist or are struggle to meet everyday lifestyle costs, but it's definitely not "insane" if this is the trade you make your living from or if you've established a lifestyle that affords disposable income for your hobbies. And for some people, $4000 for a device you have complete control over and can repurpose and tinker with to your own needs and curiosities is a much much more justifiable expense than a $200/mo rental for some narrow-access tool that somebody else controls.
- swiftcoder 21d agoIt's "insane" compared to the $1,500 it should have cost before the RAM crisis
- chews 21d agoI am pretty confident that given a $200 subscription on any of the big labs, you're getting $4000-$8000 per month in subsidized tokens... do what you wan't with your dough... and I too have a spark that I got really early (October 2025), but no, economically it does not compare to what's runnable locally in terms of quality from the frontier models. Economically, it looks like for as long as there are subscriber plans, you're better off renting. Before getting the spark, I was just using a google colab account, their $49 dollar plan allows you access to h100's and I can run qwen there in a Jupyter notebook... and if I really need that web front end I can just use cloudeflair/tailscale/the local ssh client to reverse tunnel it.
- overgard 21d agoThe cloud stuff is definitely a much better economic value, but I would argue: 1. You learn a lot more running this stuff yourself (especially since you can poke at its internals if you're interested or watch the reasoning chain.) Just being a consumer of this stuff doesn't really teach you much about it other than model & harness specific tricks that become obsolete pretty quickly. (IE, your Claude.md from 6 months ago probably needs a rewrite). Which is fine, I don't think you're going to be "left behind" if you're not a hardcore AI enthusiast or anything (I'm not), but as a guy that's always been interested in computer science I want to see how it ticks. 2. You can't really depend on this subsidization lasting forever IMO. I know the financials thing has been beaten to death but I guess I'm in the camp that it's good to be in control of your tools so that you can go elsewhere if the economics change. I like to check in with ccusage pretty frequently, and honestly like if I were paying API prices for Claude I'd probably be paying thousands a month.
- booty 21d agoCompared to pricing from 3 years ago, it's insane. The Sparks admittedly are kind of anemic: 273GB/sec is the same bandwidth as a midrange 4060, although (depending on how you configure things) you can effectively have much greater bandwidth by connecting them. Compared to 1-2 years worth of LLM tokens for a full-time software engineer making $100K+/year, a one-time spend of $12K for 4 Sparks for on-prem private LLM inference starts looking reasonable, particularly if privacy is an important consideration. It starts looking even more reasonable if running something like a private cloud to service multiple developers because then you likely need less hardware per developer. (Also, it is going to be a long time until RAM+GPU prices return to what we used to call "normal." If ever. I am not endorsing the current state of affairs and I am not saying you wrong to find it insane, but it is definitely the new reality)
- TacticalCoder 21d ago> $4,000 isn't priced insanely? ye gads It depends. My bicycle was in the 5-digits brand new (now I paid it 1/5th of that and I do thank the first owner for that: the 8 000 out of 10 K I saved were put into stocks, that's his opportunity cost, not mine). Or I know a great many a going to cry "audiofool", but I can say with certainty the following does sound better than the stereo setup of those crying audiofool: https://youtu.be/TQg9FTBMcTQ https://youtu.be/TQg9FTBMcTQ (not my setup but I've got those speakers: same thing, 15 K EUR brand new for the pair... Previous owner forked the money to buy these brand new and, well, I didn't... And I just hooked them to a wonderful, cheap, fully-integrated Yamaha amp: amazing sound). If your hobby is DIY job around the house, the cost of tools can very quickly add up too: having 20 K worth of tools is definitely not unthinkable. You like old cars? Pricey hobby. Some here even track their cars: tires and brake pads budget (and overall car budget and depreciation)... Through the roof. There's a saying that you're not really into computers if your setup doesn't cost more than your car. Is $16 K ($4 K x 4) a lot? It's six months of rent for me and for many here I'm sure. It's not "crazy crazy". Can anyone afford that? Definitely not. But there are way more insane things out there. And thanks to the individuals that go through to all the pain of setting those up, we've got feedback, tutorials, explanation, numbers, etc. as to how to run those at home. For example I helped my brother set up VMs and GPU passthrough and he's now running uncensored models locally and showing me the different answers between the uncensored models and the commercial, censored, ones. So to GP who bought four of these: we need more people like you on HN, keep it going, blog about it, be "crazy"!
- Tepix 21d agoThey only went up from 3000€ to 4000€ which isn't a lot. For comparison the cheapest Strix Halo 128GB went from 1600€ to 2600€ in the same timeframe.
- dgellow 20d agoSo, literally the same 1k increase?
- Tepix 18d agoYes! But, you know, percentages...
- swiftcoder 20d ago> They only went up from 3000€ to 4000€ which isn't a lot. Keep in mind that the DGX Spark was delayed quite a while, meaning that starting 3k price tag is already well into the RAM crisis - just 2 years ago, a 128GB DDR5 kit could be had for $600
- bmurphy1976 21d ago~$4000 USD each on Amazon, $175 for the cable.
- mmastrac 21d agoThe cables are ~USD $50 from AliExpress although I'm not sure what the tariff situation is for Americans (I think I paid $75 all-in CAD for them)
- nijave 21d agoBig "depends". Ali has "ships from China" and "ships from US" stuff. In a lot of cases, though, Amazon also tends to have cheap Chinese knockoffs for a comparable price (although sometimes their product ranking buries them and/or promotes the more expensive knockoffs) Edit: Yeah I see an ONTi QSFP56 on Amazon for $45, 10Gtek QSFP112 for $62
- cmrdporcupine 21d agoI mean, I have the same machine and the pricing is only what it is because it has that 1TB nVME in it instead of larger. nVME prices are insane and have been for months. Reality is on a single spark I'm constantly running out of room and it being an odd size M.2 slot it's a pain to upgrade. I'm setting up a NAS over RDMA via ConnectX though, that's fun.
- mmastrac 21d agoI already have a recycled QNAP-now-TrueNAS that has 10G connections into the fleet so the 1TB doesn't bother me at all. I did some rough math and I don't think I'll ever need to load weights off NFS for what I'm doing so far, but the capability is there.
- ycui7 21d agobecause it has 1T ssd not 4T
- niutech 18d agoAlternatively, you can buy cheap modded RTX 3080 20GB VRAM for $500 on Alibaba and pair 2 of them.
- 0xbadcafebee 21d agoIf you used the bare API pricing, 1M tokens @ 30% input/70% output/50% cached, you'd pay $0.05805. Even with four discounted sparks, how much are you paying for the same tokens/distribution?
- swatcoder 21d agoThere's soooo much by way of experiments, explorations, tinkering, and even projects that you can't possibly pursue through a some SaaS API. The more reasonable comparison is against rented GPU's, while looking at tradeoffs in latency and upload/download/storage/instance management overhead. Buying hardware for local models is meeting a wholly different need than buying tokens through OpenRouter or whatever.
- mmastrac 21d agoFor me it's entirely because I have a bunch of projects with my own personal data that would be tough to do with openrouter/claude or any other cloud. For example, I have a small posix-shell-based LLM harness that can SSH into my NAS and run organization tasks using the local DS4Flash that I have right now. It's already been a massive help for me to keep me organized, and that's just 2x DGX Spark's worth of compute.
- vehemenz 21d agoIt cuts both ways. A GPU in your basement is a depreciating asset with fixed computing power and consumes electricity. Switching model providers is trivial.
- bityard 21d ago> A GPU in your basement is a depreciating asset All decades prior and up to about a year ago, I would have agreed with you. My Framework Desktop, however has appreciated in value by 75% since I bought it. Will it stay there for a long time? Probably not. But it shows that there are no hard and fast rules about things anymore.
- Aurornis 21d ago> I would love to see an Opus-4.8-level local model but TBH I just haven't got there yet. The models I've tried so far _are_ good but they aren't able to solve tough technical challenges, regardless of harness/prompting/etc. Agree. It doesn’t even have to be local, using models in this size class through OpenRouter will reveal their limits if you work side by side with Opus level models regularly. There are a lot of social media posts about people cancelling their Anthropic or ChatGPT subscriptions after installing a local LLM. I’ve used local LLMs a lot and I spend a lot of time with frontier models and the difference is still huge. As far as I can tell, the social media posts about local LLMs replacing frontier models are either wishful thinking, engagement bait, or people who must be working on much simpler projects with a much higher tolerance for slop than I have.
- disiplus 21d agoTo be fair, there is no 3 turns that I don't have to jump in into what Opus 5 is doing. There is either some regression or my prompting skills are so much worse now. Flash is not perfect and honestly some things depend on how big context do you keep. So I'm keeping like a really short context with my flash, but it works okay, even though it has a tendency to overthink, and yeah, I run it always in max effort mode.
- Implicated 21d agoUse Opus 4.8. 5 is absolute garbage. Don't use DS4 Flash in max effort mode. It's just spinning its wheels, in my experience (I have a harness for testing models with 25 real bugs/features/etc from my real projects that I measure outcomes against) DS4 flash does _worse_ with max effort. It will literally have the right approach and reason itself away from it.
- disiplus 21d agoi have heard that for max effort in flash and it can be true, but overall it still performs better then the high, i run a mixed q2q4 quant.
- 21d ago
- disiplus 21d agoI will give it a try, but from the benchmarks it never exceeds the DS4 flash benchmarks by significant margin and And I feel that the throughput that you will get on those machines or what I'm getting with my local hosted flash will be so much worse that it's not worth it.
- wolttam 21d agoHopefully you also bought a switch
- whalesalad 21d agothey have 2 interfaces each so you typically daisy chain them
- wolttam 21d agoThat will hurt latency and latency is very important for good tensor-parallelism performance
- mmastrac 21d agoI think I need one now that I have four - reduce is ring-oriented and still works I believe, but IIRC you only get 200gbps if you use _one_ of two connectx ports.
- deleted 19d ago[deleted]
- metadat 21d agoQwen 3.8 27B is around Opus 4.8 level of capability on the Agentic Intelligence Index (52 vs 57). In my testing the locally hosted Qwen is good enough that looking at a given piece of work output I couldn't tell you which model was behind it. https://artificialanalysis.ai/models/qwen3-8-27b?models=gpt-5-6-luna%2Cclaude-opus-5%2Cglm-5-3%2Cgpt-5-6-sol%2Ckimi-k3%2Cqwen3-8-27b%2Cclaude-opus-4-8 https://artificialanalysis.ai/models/qwen3-8-27b?models=gpt-...
- Implicated 21d agoAs a counter to that - I've tried various flavors/quants/full weights and Qwen 3.8 27B has been entirely useless at anything non-trivial. Sure - it can do some boilerplate work (though, even armed with a well written spec and working within a very well known framework it went off the rails and did things in a way that were... um... questionable at best) but I don't see it as anything more than a personal assistant style model. Zero chance I'd "work" with it, I spent days trying to get it to do something for me that was usable that I didn't have to have reviewed and refined by a frontier level model or myself. Couldn't do it. The idea that qwen 3.8 27b is _anywhere near_ Opus 4.8 is laughable. Pure benchmaxxing. DS4 Flash 0731, on the other hand, wildly opposite experience. Would recommend. GLM 5.2 - even quanted down to a hybrid 4/3 bit setup is amazing for everything but the hardest/most complex stuff in the same projects/realm.
- skohan 21d agoI've had the exact opposite experience. I've been using 3.8 for my daily driver since last week, and I've gradually been giving it more and more complex tasks as it continues to deliver high quality results. Now I am basically handing off large complex features, and 3.8 is doing the planning, task breakdown, implementation and review with just a few notes from my side. The tradeoff is time (especially on RDMA4 hardware) - it does take a long time and spend a lot of tokens to get to the result, but I've found I can trust the results enough that I can queue a lot of work, essentially have it running all the time and achieve a decent velocity. It's the first small local model I've felt like I can do real work with.
- comandillos 21d agoI am surprised. I've been using DS4 Flash (0731) for weeks now and it works perfectly fine as a replacement for Claude in a large variety of cases. It requires a few more iterations, sure, but it's useful enough to not need a Claude subscription anymore. Among the things I do I've been reverse engineering, writing complex C++ code...
- badatnames 21d agoDS4 Flash absolutely kicks ass for reverse engineering and bug hunting. Almost no point in considering paying for a bigger model, although it's possible the stuff I've fed it (wide variety of older DOS/Windows stuff and device firmwares) might be easier targets.
- comandillos 21d agoI've been reverse engineering LEON3-FT SPARC v8 BE code, so I wouldn't say it's common :D. When attached to Ghidra through a MCP the things you can do with this are simply crazy.
- badatnames 21d agoCurious what MCP setup you use? I'm not sure which one I have wired up, but I have to restart Ghidra every time I change file. I think it's either LaurieWired's original or a fork of it
- comandillos 21d agohttps://github.com/bethington/ghidra-mcp https://github.com/bethington/ghidra-mcp . Works flawlessly.
- badatnames 21d agoHa - I saw this, but took one look at the slopfest README and it sorta scared me away. Will give it a go thanks.
- djfobbz 21d agoHow much did you drop on these 4 sparks?
- mmastrac 21d agoSparks + cables + 10g SFP+ came out to ~$21,500 CAD
- satvikpendem 21d agoSparks don't have enough memory bandwidth, for the same 20k you're better off buying RTX or Apple M5 Ultra machines.
- zackify 16d agoI ran 30M tokens through for 50c... insanity that this is possible. and it really is opus 4.8 level.