8 ms·
That's not even remotely close to being true, even once you account for capex. You have to look at the actual usage, look at the token limits. Even if you're pa
by ux266478 15d ago
That's not even remotely close to being true, even once you account for capex. You have to look at the actual usage, look at the token limits. Even if you're paying Anthropic $200k/month for scale-tier, you're going to blow through your token limits trying to run max output 24/7. Three users running Opus 4.8 at max non-stop will probably clean your monthly allowance from daddy Dario in less than a week.
With an 8x MI355x cluster at full tilt and including cooling, your power draw runs ~17kW. That's what it looks like when it's running full tilt. To be fair, hey that's pretty expensive. It does mean 8 multi-trillion parameter models unquantized running 24/7 without pause. And you get the full month like that, your monthly token limit is the time in a month. That cluster, the electrical upgrade, the cooling setup, and the electricity to run it all costs less in 2 months than your maximum affordance from Anthropic does in the same time period. Two billing cycles, and realistically it's more like two weeks. In 4 quarters you've wasted over a million. Like, what are we talking about here?
Now if you aren't using AI all that much, which is perfectly valid, and especially if you aren't using it at its absolute maximum, the story changes. Because even though at that point you're not paying nearly as much in electricity to run the cluster anymore, you still have the $300k+ capex to get the setup in the first place. But if we're not redlining it non-stop, then we're not really talking about performance anymore, are we? If your org never comes close to hitting token limits, it's probably because AI is rather marginal for you. Which again, is perfectly valid. I don't even use AI professionally.
Fact of the matter is, if your corp can justify the capex for a cluster and makes heavy use of AI, you are literally burning money by not having one in your building. The numbers are painfully obvious. Even deepseek isn't as cheap. This is before we get into things like LoRAs, custom inference pipelines, etc. which you know are kind of important if you actually care about model performance.
- Aurornis 15d ago> With an 8x MI355x cluster at full tilt and including cooling, your power draw runs ~17kW. That's what it looks like when it's running full tilt. To be fair, hey that's pretty expensive. Pretty expensive is an understatement. You couldn’t buy one of these if you wanted to right now. If you could it would be multiple hundreds of thousands of dollars. > It does mean 8 multi-trillion parameter models unquantized running 24/7 without pause You can’t even run one unquantized multi-trillion parameter (>=2T) model on 8 x MI355x with enough context for concurrent users. I don’t know how you think it’s going to run 8 of them at the same time. Did you mean 8 concurrent sessions? Your math is way off across this post. If replacing an Anthropic subscription for a whole company was as easy as buying a box for the office and then breaking even in 2 months, it wouldn’t be some little secret that we only discover in a comment online.
- ericd 15d ago>You couldn’t buy one of these if you wanted to right now. You can: https://www.exxactcorp.com/Exxact-TS4-149591758-E149591758 https://www.exxactcorp.com/Exxact-TS4-149591758-E149591758 . You can get thousands of tps of GLM 5.3 output out of this thing, which grades around Opus 4.8. Payoff is around 1 year vs. spot prices on these GPUs, including power.
- CamperBob2 15d agoI can't tell from the ad -- it says "supports" 8x MI350X GPUs, but does that mean "includes" 8x MI350X GPUs? For $300K I'd certainly hope so, but I'm assuming not. A system with 4x RTX 6000s costs about $60K these days, and can (as you note) trade blows with Opus 4.8 if not Fable. In fact, it'll give you a better pelican than Fable 5.1, and in less time.
- ericd 15d agoHa fair, I'd definitely confirm with a salesperson before wiring them $300k. But most of the signs on the configurator seem to point to it including the GPUs? Not going to make 30k BTUs/hr of heat without the 8kw of GPUs.
- Aurornis 15d ago> trade blows with Opus 4.8 if not Fable. Okay I love the open models, but the hype is getting ridiculous. The models you can run on 4 X RTX6000 are not Fable level.
- CamperBob2 15d agoWell, they are if you're into animating pelicans. :-P But yes, in the general case Opus is a better match. And Opus is no slouch. I'm satisfied that GLM 5.3 is just as strong as Opus. Z.AI has promised/bragged that they will be at Fable 5.0 level by the end of the year or early next year, and I don't see any reason to doubt them.
- ux266478 15d ago
- pcarolan 15d agoHere’s an experiment: purchase an anthropic pro max subscription for $200/m. Now go buy the hardware to run DeepSeek’s equivalent. In a year, who spent more?
- ericd 15d agoThat's not apples to apples on almost any dimension.
- egeozcan 15d agoIn normal times in which hardware used to depreciate (lately that's not the case and HW even appreciates, but let's not get distracted), if you calculate only with depreciation costs, plus the fact that when you have such a setup, it'd take many 200$ subs to cover your lack of limits in the other, I think it'd not be a clear victory for any side. If you just ask "who spent more in the first year" (100% depreciation) then even with 5-6 max accounts, buying HW will be a couple of times more expensive. But when does it make sense to ask that question? Maybe the SotA models will need better hardware so your investment will not be useful after a year or you'd need very expensive upgrades? But then (as in Fable case) subscribers need to spend more too.
- srcreigh 15d agoIt’s not so clear after 5 years that you’ll come out ahead. You’ll have spent $20k. The apple computer owner will probably be running local models that are better than today’s frontier on the same hardware. Idk where you live, but where I am running the M5 Ultra Mac Studio at max rated power 24/7 for a month costs C$42. The considerations against Apple hardware are 1) hardware advancements 2) early access to the best models. But it’s really not that clear. (The other guy who thought hosted models on openrouter are cheap has spent $100k in 5 years.)
- SXX 15d ago> The apple computer owner will probably be running local models that are better than today’s frontier on the same hardware. Hardware is not magically getting more memory or bandwidth. Believing there will be some magical optimizations to compensate for it is just dellusion.
- ericd 15d agoWhat're you using that monster for?