7 ms·
I don't think you're going to get many "true" answers to this. The opportunity cost of not using the latest and best models is just too much right now. Every m
by codinhood 3mo ago
I don't think you're going to get many "true" answers to this. The opportunity cost of not using the latest and best models is just too much right now.
Every month I research this and come to the same conclusion: the time, effort, and cost required to get local models (and the coding tools around them) to perform even close to Claude Code with sonnet/opus just not worth it right now. If it was, it would be distributive enough to be in the news.
Not that I'm discounting someone hasn't already solved this, just trying to Occam razor my way out of diving too deep down rabbit holes.
- jrm4 3mo agoBut you're pretty much measuring opportunity cost in tokens per second, no? I think it strongly remains to be seen whether e.g. tokens per second (multiplied or whatever by percieved quality of private model) actually means "better or more useful output." I strongly suspect it does not. (though I also strongly suspect this will be very difficult to measure because the incentive to lie about metrics here will be so strong.)
- codinhood 3mo agoIf you’re arguing that model metrics don’t necessarily translate into useful output, I agree. That’s not how I measure the success of a mode and not really the point I'm trying to make. I try to set things up and test it on my actual projects. What I’m saying is that if local models were actually comparable to Claude Code in practice, we wouldn’t be having threads like this. It would be obvious to the people using them, and it would be massively disruptive. Why would individuals and companies pay hundreds or thousands for Claude Code if they could run something locally and consistently get similar results? Every month I revisit the local ecosystem hoping the answer has changed. So far, my experience has been that it hasn’t.
- jrm4 3mo agoHaving, e.g. seen Microsoft maintain a monopoly for well over a decade, there's nothing in my experience that suggests that "quality always beats hype" is remotely true. It's entirely possible Claude is just winning the hype game.
- robertlagrant 3mo agoMicrosoft have not maintained a monopoly on search, mobile, or maps, and they seem to mostly maintain their large market segments based on familiarity, not hype.
- jrm4 3mo ago? I was speaking historically, not now
- robertlagrant 3mo agoI don't know why you're surprised; you didn't specify you meant historically.
- jrm4 3mo agoThat was the point of the "e.g."
- robertlagrant 3mo agoE.g. doesn't mean historically.
- Rastonbury 3mo agoI think they are referring to the opportunity cost of time saved on doing things a local model cannot do or fixing it's mistakes against the cost of a subscription
- MadrasThorn 3mo agoIt's great at accelerating hardware innovation however.
- sakopov 3mo agoThis seems to be the answer. Building a rig with a decent graphics card will cost $2k+ and will produce sub-par results. Might as well milk the $100/m Claude sub until open-source alternatives reach parity with today's frontier models.
- pyeri 3mo agoAt some point, there will come a saturation point for that "Opportunity cost FOMO train ride", and I think we are already past that point. Mythos class models are a whole different beasts and cutting edge on reasoning but not much use for the problem domains most developers are trying to solve. The present Sonnet/Opus versions (~4.8) will likely be what everyone in the enterprise might end up using eventually. And even though local models aren't there yet, there are budget alternatives from the families of DeepSeek, Kimi, GPT, MiniMax, etc. available through APIs of NVidida, OpenRouter, Groq, etc. which are very much Sonnet grade.
- codinhood 3mo agoYeah this is exactly what I'm waiting for. Personally, I don't think we're at that point yet. While I do think model improvement is starting to plateau (reaching a local ceiling), I'm not convinced local models are as good as sonnet/opus yet. The gap is still too much. But I'm excited for those models to reach those levels.
- mark_l_watson 3mo agoSounds like a correct conclusion to me also. I am trying to transition to a layered system: local, then OpenCode with commercial vendor APIs for models like DeepSeek v4 flash, then DeepSeek v4 Pro. With a layered approach we can slowly shift to running more locally and still get required work done. Really, my local setup is so much better than it was 2 months ago, and extremely better than 6 months ago - on the same hardware.
- gunapologist99 3mo agoRather than Occam, consider Pareto? If you truly believe that it WILL get there within the next couple of years, then you might as well start playing with it now (and, yes, you will be very surprised, especially for shorter/smaller projects or nicely modularized larger projects)
- phyzix5761 3mo agoThe opportunity cost to who? Its getting super expensive for businesses and engineers across the board to pay for frontier models.
- Gigachad 3mo agoThe cost of the hardware to run local models is still massively more expensive than the subscriptions while offering worse models. Eventually I think it will even out but right now the hosted stuff is very subsidised.
- kristopolous 3mo agothat's super contextually dependent. I use them just as essentially a decompress of what I already know that I'm doing. I legitimately use 4B models just fine. I've got a large number of tools that make this entirely feasible and a daily driver for me (like https://github.com/day50-dev/llm-manpage-tool https://github.com/day50-dev/llm-manpage-tool) ... It's not really a bitter lesson here, I can scale those 4B models easier than someone can scale their 1000B models.
- bob1029 3mo agoI've got a machine in a corner collecting dust that cost me $12k to build 2 years ago. It runs fine but it's wildly impractical to use as a daily driver (loud/hot). I keep it as a reminder to not do this again. At my current pace it would take me until sometime late 2030 to spend the same amount in gpt5.5 tokens.
- anonzzzies 3mo agoYou forget that, especially on HN, many people are scaremongering that prices will soon skyrocket. Then it will be another story... I easily run $4k+/mo on my claude sub; if I would have to pay that, I definitely would spend 12k on hardware instead and accept a dumber helper.
- jonfw 3mo agoYou are not stuck between public API pricing for frontier models via Claude and self hosted. Remember that there are other LLM providers, open models, and previous gen models, that are way cheaper that frontier Claude and still way better than what can realistically run locally
- NamlchakKhandro 3mo agoThinking Claude is leading edge... I really think you need to re-evaluate what you research you think you're doing. Claude Code is not Claude Opus/Sonnet/Haiku.
- reassess_blind 3mo agoWhat is leading edge then?