7 ms·
Open weight models from Chinese labs tend to be significantly cheaper. I think theyre absolutely needed. I can't afford 200 USD a month for personal use of cod
by jerojero 3mo ago
Open weight models from Chinese labs tend to be significantly cheaper.
I think theyre absolutely needed. I can't afford 200 USD a month for personal use of coding AI, and I don't think such prices are reasonable for most of the world economy anyway. Not to mention US firms might be giving their employees a lot more than that.
It's increasingly feeling, to me, that theres a gap building up between haves and have nots. But then, we get news of these open weight models that are reasonably priced in inference with reasonable capabilities. Yes, they take maybe 6-9 months to get there, tbh, that's not a bad trade off at all.
- narrator 3mo agoThe tokens cost the same everywhere on earth. This does hurt some cost advantages of outsourcing when tokens start to become a bigger part of development costs.
- cameldrv 3mo agoYes, but you’re paying with your data unless you’re hosting with a provider you trust or self-hosting.
- sixothree 3mo agoMy first instinct has been - well this is an open source project, what does it matter. But even then, I am guessing that using their service even for open source projects still provides them some value.
- cookiengineer 3mo agoKind of funny that you're assuming that you are not paying with your data in both cases. Do I need to remind you how LLMs are being trained? ...or that Anthropic claimed their codebase is 100% vibecoded, making it uncopyrightable by their own logic? ...or that Anthropic took down all Claude Code leaks they could find using DMCA takedown notices? ...or how do you think the caching mechanisms work when there's allegedly no data stored to be able to cache it? I'm just saying. Anything you build with online models is their training data anyways. Assuming otherwise is pretty stupid at this point.
- ttoinou 3mo ago200 is much less than the value you’re supposed to get out of it. If it’s not then yeah go ahead and use cheaper models with worst quality
- Dayshine 3mo agoI'm not sure how I'm supposed to get $200 of value out of personal use!
- LPisGood 3mo agoNote that 200 dollars of value is different than 200 dollars of profit.
- devmor 3mo agoI personally don’t find it that useful for most tasks, but if say, you get paid $50/hr for your work and it saves you more than 4 hours of work in a month, there you go.
- deleted 3mo ago[deleted]
- selcuka 3mo agoObviously this assumes that you can find 4+ extra hours of $50/hr work every month, or you can work 4 hours less. Neither of these assumptions is correct for people who work for a fixed salary.
- devmor 3mo agoThat doesn’t change value. It’s value whether or not you can maintain a profit over it.
- selcuka 3mo agoThat's the definition of "value" in a broad, economics sense but I don't think it applies to the parent comment of this thread: > I'm not sure how I'm supposed to get $200 of value out of personal use!
- tacomagick 3mo agoDeepSeek through their own API has saved me tons of tokens honestly. Even though it is not as smart as Kimi or Claude, their level of entry is very low with a top up of 2$ and Pay as you go compared to the subscription of Claude or 20$ top up of Kimi
- praveer13 3mo agoFor personal use I’m considering using the frontier models from openai or anthropic to create a plan with research and brainstorming etc with enough details for cheap models to be able to follow (glm, deepseek etc) - with openrouter - will monitor how cheap and effective that turns out to be.
- ImaCake 3mo agoYou should try out the cheaper models first. I find Deepseek v4 models pretty comparable to sonnet 4.6 but at a fraction of the cost. You might find you just don't need to use the American models at all.
- tacomagick 3mo agoFor my case Openrouter breaks Deepseek caching and charges me multiple times over what I pay for Deepseek's API, with 2$ I was able to get around 120M tokens from deepseek easily when Openrouter could only barely do 250k
- jabroni_salad 3mo agodeepseek's direct API is super loosey goosey about caching. On multiple occasions I have gotten cache hits resuming a session from the previous day.
- lionkor 3mo agoSeconding the recommendation to use Deepseek directly via the API. I've burnt 287 million tokens in the last couple of days, costing me a whopping $5.77 USD.
- 3mo ago
- Fr0styMatt88 3mo agoIf we can agree that the AI model is at least as capable as a junior engineer or new contractor, how’s that different to saying “software engineering isn’t worth $200 a month”? Has a very race-to-the-bottom feel to it. Though in the grand scheme of it, $200/mo probably isn’t the real price either. Also looking at it not just in a vacuum - paying for a product that can change what you get from under you doesn’t seem great anyway. At least with a locally-hosted model you know what you’re getting.
- matheusmoreira 3mo agoYeah. There's no way to verify what these providers are doing. The real future is running these models at home. Opus level inference on our own hardware would be a dream come true.
- IncreasePosts 3mo agoHow will anyone running home instances be able to compete against people paying some money running much more powerful models on much more powerful hardware?
- Fr0styMatt88 3mo agoIt’ll be interesting. I’m using Qwen3.6:27B at home and mostly Sonnet/Opus (depending on the complexity of the task) at work. You have to break things down into smaller chunks for the local models. For the bigger cloud ones they can do a lot of the broader thinking.
- fragmede 3mo agoTime is money, but apparently now thinking is money as well. How much is it going to cost to think harder? If it's, say, $10 to use a bigger cloud model, it becomes easier to qualify the cost of thinking.
- jimbokun 3mo agoAt some point it will be hard for us to tell the difference.
- ImaCake 3mo agoSignificantly cheaper than comparable models if you are using openrouter [0]. Just yesterday I spent roughly 13 cents centering some divs using Deepseek in a personal project. It would have been north of $1 to do that with a US frontier model. 0. https://openrouter.ai/compare/z-ai/glm-5.2/anthropic/claude-opus-4.8 https://openrouter.ai/compare/z-ai/glm-5.2/anthropic/claude-...
- ipaddr 3mo agoFor centering divs the free models opencode offers can easily handle that work. DeepSeek V4 Flash is pretty decent.
- ImaCake 3mo agoSure, but something that is “sonnet tier” is going to get there faster and with less pain. Well worth the 13 cents!
- ipaddr 3mo agoFlash will get their faster then the sonnet tier which involves reasoning which is slow. And you don't need reasoning to center divs. The sonnet tier sits below claude or chatgpt in terms of price but costs so much more than free models. If you are breaking downtasks now I'm not sure that 13 cents is worth it.
- arikrahman 3mo agoSomeone else on this forum put it well, U.S. is trying to achieve AGI at all costs, while Chinese models are seeking widespread adoption.
- azinman2 3mo agoI don't think anthropic/openai/google aren't also seeing widespread adoption. In fact they already have they already have the marketshare.
- Turskarama 3mo agoThe difference is that the US companies are using it as a means to an end, they need to make just enough profit that the investors don't all get cold feet before they get to AGI. The Chinese companies on the other hand are trying to be profitable immediately, which means that they're going slower to save development costs.
- lionkor 3mo agoNone of the AI companies in the US are on the path to AGI. They are, however, on the path to claiming they have AGI, then subsequently not releasing it and only giving it to the US government to make drones that can bomb the homes of political dissidents.
- dotancohen 3mo agoWhat kind of off topic political ideology spam is this? Do you not think that the Chinese kill their enemies? The Chinese are genociding Uyghurs as we speak, purely for being Muslim, in numbers that dwarf any harm the US has done.
- lionkor 3mo ago> in numbers that dwarf any harm the US has done. The list of wars the US is or was actively involved in[0] is SO LONG that the Wikipedia page is split into multiple different pages. The main relevant ones are 20th[1] and 21st century[2], for which you better get a good grip on your mouse to scroll down. I urge you to use your favorite AI to give you a rough summary of direct and indirect casualties of just those wars directly caused, started, or provoked by the US, from these lists. For example, the "war on terror" alone has, so far, seen around 4.5–4.6 million+ people killed, and at least 38 million people displaced. [0]: https://en.wikipedia.org/wiki/Lists_of_wars_involving_the_United_States https://en.wikipedia.org/wiki/Lists_of_wars_involving_the_Un... [1]: https://en.wikipedia.org/wiki/List_of_wars_involving_the_United_States_in_the_20th_century https://en.wikipedia.org/wiki/List_of_wars_involving_the_Uni... [2]: https://en.wikipedia.org/wiki/List_of_wars_involving_the_United_States_in_the_21st_century https://en.wikipedia.org/wiki/List_of_wars_involving_the_Uni...
- throwaway-blaze 3mo agoJust don't ask it to tell you the events of June 4, 1989.
- jampekka 3mo ago[flagged]
- girvo 3mo agoNot that it matters but most of the open weight models aren’t actually censored that way: they run another layer on top of to do that. At least some of them do, Step 3.7 Flash locally happily tells me about the Tiananmen Square massacre
- swingboy 3mo agoMy work involves asking LLMs about both Tianenmen Square and what’s going on in Gaza, so I can’t use Chinese or American models!
- matheusmoreira 3mo ago> It's increasingly feeling, to me, that theres a gap building up between haves and have nots. People speak of a permanent underclass. https://www.nytimes.com/2026/04/30/opinion/ai-labor-work-force-silicon-valley.html https://www.nytimes.com/2026/04/30/opinion/ai-labor-work-for...
- fbrncci 3mo agoYou made me realize something. I routinely spend upwards of 500$ per month on LLMs for coding (expensed towards clients). However I live in a place where 500$ is around the avg. salary. I’m lucky that I know my way around western clients. Clients who pay these expenses and are happy to work with me because I am still about 50% cheaper than local talent in EU/US, while my salary at home converts to an upper class income at the highest tax bracket. Which of course causes some unfairness on both ends. Nobody here can compete with me. I often use left over tokens on local client projects; which despite lower pay, still pays off because they now take hours not days or weeks to complete. And nobody in the local clients talent pool can compete with me; unless they charge about half the market rate. Take away my 500$ monthly grant; and I’d be more or less screwed. Better open models will more or less start to reduce this advantage. It’s not like I positioned myself here on purpose. But it’s definitely a „right place, right time“ situation.
- swader999 3mo agoIf you are running multiple agents your cost to them should be multiples less what their roi is.
- fbrncci 3mo agoMy costs are 0$ as any token or subscription spend on agents is invoiced as an expense to my clients.
- kreelman 3mo agoThanks so much for being bold enough to be fairly open about the costs, how you arrange billing and the advantages that's given you. I've been fooling around with DeepSeek 4 agentically. It's probably not as good as Anthropic offerings, but even those seem to be roiled in politics and strife and DeepSeek 4 is very good IMHO. I'll later try out GLM. I'm in Australia. The government has set up a "return and earn" scheme to keep aluminium cans, plastic bottles and paper drink cartons out of the waste stream. A laudable project. The money you make from return drink containers is pretty low, $AU 0.1 per container. I've participated to get the rubbish out of natural water streams and to make a nano amount of money on the side. When I looked at the costs of an app I was getting DeepSeek to help me with, I realised that the several hours I'd spent learning and building had cost something like 8 recycled containers. In my head after doing some DeepSeek stuff, I calculate a "cans per app" metric for myself for fun. I may even setup a simple graph to view my costs that way. I kind of hope the Anthropics of the world get enough price competition from sources like DeepSeek and GLM to drop their prices significantly. Time will tell. I'm using the Chinese DeepSeek provider, so everything done there could potentially be taken and used by the CCP... But this is hobbyist learning. There is probably a market for Deepseek/GLM served from non CCP available servers. I might even look into how hard that would be to setup here. I also hope that inference focused hardware will come to the fore, reducing energy use and cost. Realistically this will take time though, on the order of years. Here in Oz, we have community batteries that community members can charge and later draw from. Their electricity prices are competitive. I wonder if someone could setup something like a community battery to run data centres... That way reasonable environmental consideration could be given to inference power generation... This might not work in a market like the US or Europe, but small market size might be an advantage... Who knows.
- brian-armstrong 3mo agoI read these stories and I can never figure out how people are managing to use these $200 plans. If I really go full bore, I can sometimes max out the $20 plan. Even then, it already produces more code than I can reasonably review and merge.
- ipaddr 3mo agoI've maxed out my chatgpt plus the first week and that include an smf forum rewrite. Trying my best I haven't been able to max out again. Things are setup that you need to max out your 5 hour window multiple times which becomes a job in itself. At work I'm struggling to keep my claude bill around $500.
- girvo 3mo agoSimple: a lot of the people claiming they’re reviewing the output of these models are lying. Also if you run the “loops” they’re now yapping about, it will burn through enormous amounts of usage as well.
- hgomersall 3mo agoI can't even keep up with the chain of thought needed to manage a single session, let alone review. I typically never exceed 30% of a 5x plan. Fable took me almost to the limits, but not Opus. Claude design hits things harder, but still not to saturation.
- theoli 3mo agoExactly this, it’s the loops. The first 50k tokens of a task is by far the most valuable. But when left to run independently, the agent will consume millions of tokens of error messages from running tests and discovering a minor syntax error, a missing import, a method call with incorrect parameters, etc. Then it will write some helper program while debugging the main task and get into the same loop debugging minor errors in the helper. From my experience, the vast majority of tokens consumed by Claude Code on totally independent tasks are consumed fixing minor mistakes it just made.
- RugnirViking 3mo ago
- alpineman 3mo agoWith open weight models there is true inference competition. Whoever can serve the model at the lowest price. And the consumer wins. Capitalism, served by China.
- giancarlostoro 3mo agoAs much as I don't like Mark Zuckerberg, part of me wishes he would get his head in the game and compete with these models, he's literally got all the capability to do so, and he could easily sell the model through deals with GCP, AWS, and Azure. Hell, Amazon needs a hot model they can host that's exclusive to them I feel like, maybe he can work something out with them, whatever the case, it seems so glaringly obvious to me, I'm not sure why he hasn't taken a stab at competing with Claude Code or at least frontier open models and then cutting a deal with cloud providers to recoup the costs of maintaining said models. He's sitting on a frontier model letting it burn a hole in his wallet that could actually pay for itself.
- khurs 3mo agoMeta internally have been using Google Gemini "Meta has been using Google’s Gemini large language model for most of its moderation and customer support, but staff have recently been told to switch to Meta’s new foundational model, Muse Spark, the people said." https://www.ft.com/content/39251a31-4a9d-4870-b86c-dc6353d67fdd https://www.ft.com/content/39251a31-4a9d-4870-b86c-dc6353d67...
- giancarlostoro 3mo agoIt feels really insane to me that they have a model that could be better, but its just sitting there burning a hole in his wallet instead as he chases trying to recreate Grok's companion thing.