6 ms·
To me, most local models work just fine for anything you can be patient for. If I want something quicker, I will go to a SOTA model via API, but with multiple 3
by jermaustin1 20d ago
To me, most local models work just fine for anything you can be patient for. If I want something quicker, I will go to a SOTA model via API, but with multiple 3090s, I have never really needed a hosted model for a lot of my experiments.
For code, they are great, but for creativity for NPC controllers, they leave something to be desired, but work well enough for testing, so I don't burn tokens until I'm actually playing my games.
But nothing one-shots a prototype better than Fable 5. I can have a prototype built in 30 minutes, hooked up to my local LLMs and Claude Code is very good at testing the interactions and even tuning the prompts of the NPCs for better experiences.
- __float 20d ago"with multiple 3090s" is quite a bit of burying the lede for "most local models work just fine", don't you think?
- jermaustin1 20d agoHaving multiple 6 year old cards doesn't seem like it's that big of burden for local LLMs. I get that a lot of people don't have them. And a single one can be VERY performant. And the smaller models like a 7B can run on much smaller hardware like a mid-range [3|4|5]060. My entire AI Dev Box cost $4500 in parts. 128GB RAM, i7-10700, 1TB and 2TB SSD, and 2x 3090s. Today's prices and inflation have definitely made that price tag seem a lot better than it was, but it was an investment in all things GPU that were happening in 2020 (crypto, blender, image gen), then LLMs exploded.
- thayne 20d agoA single, used 3090 costs more than I have ever spent on a computer.
- 9cb14c1ec0 20d agoYes, the tunnel vision around local models on this site is crazy. The percentage of people in the world who can afford the hardware is extremely low.
- layer8 20d agoIt seems roughly similar to the pricing level of personal computers in the early eighties (i.e. IBM PC and Apple Macintosh). I’d expect prices to come down significantly over the next few years. Not so much in the next year or two, but after that.
- oblio 20d agoJust like for warships, the complexity and cost of building cutting edge hardware has grown exponentially up to a point where a significant chunk of the world's computing is dependent on 2 companies: ASML, TSMC. We shouldn't extrapolate linearly from examples from the 80s.
- layer8 20d agoNo, but I wouldn’t expect it to stagnate like with Intel in the 2010s either. Maybe the biggest caveat is that most people will be fine with using cloud providers, so the market for non-server hardware won’t be subject to as much competition.
- inigyou 20d agoThat's in nominal dollars. However, inflation since then has been about a factor of ten to fifty and it hasn't trickled down at all.
- jermaustin1 20d agoI don't think there is tunnel vision. I'm just saying that I have a couple 3090s I invested in a handful of years ago, and they are still going strong today as multiple GPU-needing technologies emerged. I'm not saying everyone has to run local LLMs, because the APIs are in a race to the bottom, and my $10 of OpenRouter credits I bought months ago is down to $8.94 because most models give you MILLIONS of tokens for a US Quarter.
- 9cb14c1ec0 20d ago> I'm just saying that I have a couple 3090s This is tunnel vision. The percentage of people who could afford the hardware you could at the time you back it so vanishingly small. I do not know a single non-tech person who has multiple graphics cards in a single computer.
- wafflemaker 20d agoMy single 3080 runs so hot I don't need to warm my room in winter, and have to play games in my underwear in summer.
- Karrot_Kream 20d agoI keep coming back to this: why do I need to run a local model on my own GPU? Open models can run in dedicated clouds and while, yeah, they may be more expensive per token than my own GPU, when accounting for depreciation, energy usage, and opportunity cost (money not spent on my GPU will instead sit in my portfolio appreciating with its particular blend of returns), I'm pretty sure I break even or even net lose money with a GPU. Don't get me wrong, there are advantages to a fully local model in that, I can have agents looping 24/7 even when my internet is not working. But this is niche enough that if I had to price the advantages they don't seem worth it. If I'm willing to pay the Openrouter tax, I can fire up Openrouter today and just get access to whatever model I want, and still pay a fraction for tokens as what I'm paying with the big guys.
- throwaway219450 20d agoUnless you value privacy, pay for openrouter. You still get the benefits of cheap tokens and programmatic usage. 3090 pricing is something of a wild card. Since the only big-mem consume cards are the xx90s, and a 5090 is pushing $5000, resale value has gone way up. The bottom hit ~$700 last year. It's still a very good GPU, if power hungry.
- zamadatix 20d agoI got a great deal on ~72 TB of NVMe right before storage prices shot up, doesn't make it any less ridiculous that I have it or any more relevant to people talking about building a NAS now. 99% of people, even in tech, do not have the stupid amounts of hardware people like us hobby on.
- sroussey 20d agowhere? i would love that.
- zamadatix 20d ago"Where'd I buy it" or "where is it now" ;)? It was a 96 core gen 4 epyc+supermicro board build with consumer NVMe drives on 1x16->4x4 "dumb" bifurcation cards. I had to get a few MCIO-> PCIe adapters as well to get the full lane coverage. Mounted in a standard EATX compatible consumer case with a consumer PSU and a lot of Noctua fans - surprisingly cool and quiet for what it is. Motherboard+CPU I got from Ebay. Rest from the best MicroCenter/Amazon/Walmart deal of that day. Bought juuuuust before the AI pricing apocalypse, largely by pure chance.
- oceanplexian 20d agoMost people in the US have a car, and the average new car is $40,000. Hell where I live a middle class consumer will spend double that on a Boat or an RV and think nothing of it. These aren’t elite tech workers. It’s not unfathomable that if a personal, generally intelligent local AI provides enough utility and doesn’t require you to tweak CLI flags millions of Americans would want one.
- spockz 20d agoSpending that kind of moment on a product that gives you personal happiness for years up to decades and then will still have residual worth, which people save up for ages for, is an entirely different proposition than buying a product that may make you faster professionally, but which in the short time can also be achieved by a few dollars worth of subscriptions to a hosted model for even greater effect.
- xnx 20d ago> 2x 3090s You could sell those and have enough money to pay for hosted inference for years.
- jjav 12d ago> You could sell those and have enough money to pay for hosted inference for years. From a quick search a 3090 looks to go for about 1500-2000 USD. So let's say $4K for two. I'm spending far over $1K/month (employer-paid) on cloud AI, so if that could be anywhere near comparable we're only looking at less than a few months break-even. Less really, because some months are more expensive. This month I'm up to ~$500 and it is only day 4 of this month.
- jermaustin1 20d agoThey cost more to run than hosted anyway. But that isn't the point of having them. They are a playground, a backup when the internet is down, or claude is down. They can render Blender scenes pretty well. They play any game I want. You can do each of those at various hosts and own nothing. Or own a couple "over priced" cards and do it all at home on battery power for a few hours while the power is out.
- robotresearcher 20d agoFor me it’s more that you can show them your financial and medical data without BigCo looking over your shoulder.
- lostmsu 19d agoBigCos you are referring to are earning their money doing state of the art research, not so much from your financial and medical data. That was the age of Internet ads, which was over since AdBlock was created for anyone concerned.
- Gecko4072 20d agoBut after all those years you’d still have 2 3090s, which are now about 6 years old and still holding value.
- vel0city 20d ago[dead]
- bitexploder 20d agoNot really. 2 years ago that was a pretty normal amount of GPU hardware for a hacker or gamer. It's all relative. They are not accessible to most people yet, but for someone that cares and is a technologist? Likely accessible.
- sroussey 20d agoI have trouble getting simple extraction to work sometimes. I have a block of text describing people and their roles at a company and their ages, and i asked for structured results of an array of these things with the text span that it appears in and all i can say is: nope.
- quotescoreai 20d ago[flagged]
- jermaustin1 19d agoI've done pretty decent local prose->json extraction using Qwen and Phi and Gemma. I'm sure most of it comes down to prompts, and all of them run over 100tps on a 3090. Smaller cards will likely be slower, but Qwen3.5 9B is small enough to fit on most consumer cards.