6 ms·
That's only half the reason it's expensive. The other reason is that it would likely take years to spend $4000 (plus the real cost of electricity) worth of tok
by vehemenz 21d ago
That's only half the reason it's expensive.
The other reason is that it would likely take years to spend $4000 (plus the real cost of electricity) worth of tokens on a 3rd-party provider that's running a similar limited, DS Flash type model. By that time, the hardware will be obsolete, assuming it's still operational.
- swiftcoder 21d ago> it would likely take years to spend $4000 (plus the real cost of electricity) Since that cluster only yields 20-30 tok/s on that size of model, at least a decade before the hardware breaks-even with current token costs, and that's not counting electricity. Assuming continued downward pressure on token prices, and the cost of electricity, it never pays for itself.
- dannyw 21d agoAs a counterpoint, my homelab/home-LLM hardware has appreciated in value by about 60% since I bought it. Of course, it's not real unless I sell, and the value will eventually go down, but so far I have significant paper profits. Also, DeepSeek token prices are continuing to _increase_, not decrease.
- swiftcoder 21d ago> DeepSeek token prices are continuing to _increase_ One increase does not a trend make. And the current crop of models are now undercutting deepseek flash...
- Implicated 21d agoYou can't possibly think that it's going to get cheaper and cheaper to pay for tokens though. Right? Have you seen what's happening with Codex/Claude subscriptions? Deepseek raising API prices.. We've been getting subsidized tokens for some time now and as the hardware costs skyrocket these labs/people with inference compute are going to continue to clamp down.
- keheai_harvey 15d ago[flagged]
- swiftcoder 21d ago> You can't possibly think that it's going to get cheaper and cheaper to pay for tokens though. Right? Absolutely I do. Each generation of open-weight models has come with significant efficiency improvements, and there are significant hardware gains on the horizon: both increasing competition from Chinese chip manufacturers, and new custom silicon from the established players. And unlike Anthropic and OpenAI, most of the pure inference providers aren't massively leveraged - the more hardware they can bring online, the cheaper they can serve tokens.
- maxglute 21d ago$40,000 GPU is like few pennies in sand. Only mildly hyperbolic. But a GPU fresh out of fab is $2000 after ASML, TSMC and inputs get their 50-75% margin, then somehow $40k laundered through US financialization / Nvidia margins. Commoditized GPUs shouldn't cost more than 1-2% current price once there's competition.
- rsdhrghrdh 21d ago[flagged]
- maxglute 21d agoCompelling argument from only msg on new account.
- selectodude 21d agoWe’re getting subsidized training. The inference is not a loss leader. And since providers can run hardware at 100 percent 24/7 their per token cost is going to be far below mine, regardless of how long I’m willing to wait for a token to come out.
- nkozyra 21d agoI don't understand how people don't consider this. Plus you're spec'd out of near-SOTA level in months. The only reasons to actually do this are a) you have a lot of dispensable income and are a hobbyist/tinkerer, b) you have real, legitimate privacy concerns or, relatedly, c) you're doing something you don't want to get flagged
- billiam 21d ago"you don't want to get flagged" Ding!
- swiftcoder 21d agoWhat is actually getting you flagged by the openweights inference providers? Thus far I haven't hit any of the reverse engineering or infosec guardrails that Anthropic is so keen on
- nkozyra 21d agoWhile I'm sure some of the open weight providers do this as well, I think the comparison is frontier labs v local inference.
- swiftcoder 21d agoI'm not sure that is the comparison - the OP is planning to run an open weights model locally, the obvious comparison would be paying a hosted provider to run the exact same model
- nightski 21d agoNot everything is about pure cost. Maybe I don't want to sell my soul supporting the frontier labs because they are straight up pure evil?
- nkozyra 21d agoI barely see a difference between buying the hardware that feeds (and often colludes with) those labs, at least not as a moral stance. Even if you trained your own model, you'd be committing some of the same sins, paying for the same hardware that drove it, etc. But if you're using some open model, you're standing on the shoulders of the same corrupt giants. I feel like when people say this is due to moral reasons, it's to justify an expensive hobby.
- girvo 21d ago> it never pays for itself. Exactly; its a development box for fiddling with GPU hardware with a large amount of video-addressable memory. It's not an inference box, really, though it's neat that I can at all!
- KronisLV 21d agoBut if you can use cloud models, why wouldn’t you use SOTA? For 2400 USD or less per year you can get pretty huge amounts of benefit out of that (though at the whim of whoever you are giving the money to).
- mandeepj 21d ago> The other reason is that it would likely take years to spend $4000 (plus the real cost of electricity) worth of tokens on a 3rd-party provider That's just a one-dimensional thought! Your own hardware gives you complete control, and it doesn't time you out for 4 hours, unlike those vendors.