5 ms·
I'd be more OK with it if prices weren't subsidized so much and people actually had to pay to ask opus how to pee, then maybe we would realize we don't need bee
by amarcheschi 2mo ago
I'd be more OK with it if prices weren't subsidized so much and people actually had to pay to ask opus how to pee, then maybe we would realize we don't need beefier models for everything
- simianwords 2mo agoThe fact is, you wouldn’t be okay with the real prices as well. All indications point to opus not being subsidised but having huge margins. GLM 5.2 is not subsidised - it is an Opus tier model that costs a small fraction. I doubt you would be okay and all problems would suddenly vanish.
- ekidd 2mo agoGLM 5.2 isn't quite modern Opus tier, as seen in this comparison where Opus 4.5 scores 4/5 on some coding tasks where GLM 5.2 scores 0/5: https://www.tryai.dev/blog/gpt-5.6-build-off-12-models https://www.tryai.dev/blog/gpt-5.6-build-off-12-models But yes, GLM 5.2 is cheap. But the real standout on price is DeepSeek V4 Flash, which competes, more or less, with models in between Sonnet and Haiku. From third-party providers, it costs around $0.09/M, $0.18/M out, compared to $3M/in, $15M/out for Sonnet and $1M/in, $5M/out for Haiku. To get the price of DSv4 (Flash and Pro) so low, DeepSeek did a lot of innovative optimization work that will likely show up in other open weight models in the future.
- ronsor 2mo agoI think GLM's problem is a lack of vision input.
- trollbridge 2mo agoThere isn't much evidence at all that inference is "subsidised" (and by whom?) Training is quite expensive and it does look likely that the American providers have been doing that at a loss. In any case, you can go buy a MacBook Pro M5 48GB or an AMD R9700 and run Qwen 3.6 35B-A3B (a very capable model) and the only "subsidy" is you plugging it in, and 140W is not exactly a huge amount of power (roughly 50¢ per day if you run it 24/7 at 100% load, which it is very unlikely you will).
- amarcheschi 2mo agoI agree with using smaller models, it's just that the majority of people I know feel like they need the biggest, beefier, behemoth model possible (with the longest thought setting) and consume much more than necessary when a flash or smaller model would be OK. I would also like to be able to use a smaller model, but given ram prices I would have to sell a kidney to buy ram now
- trollbridge 2mo agoMost people who use a free or $20 a month plan are already using smaller models, and the mainstream chatbot services will route requests to a smaller model often without really telling the end user. You can run Qwen-3.6 on a 32GB card which will set you back about $1400, or $400 of just RAM if you want to run it on a CPU.
- mschuster91 2mo ago> There isn't much evidence at all that inference is "subsidised" (and by whom?) Well... why else would the major providers now tighten the screws on per-token pricing?
- blfr 2mo agoBecause they thought they could. Turns out the Chinese and Elon had other plans.
- mschuster91 2mo agoThe Chinese providers are just as much getting subsidies from the CCP, and Musk/SpaceX is (indirectly) raiding retirement funds to fund the bonanza.
- blfr 2mo agoThe truly cheap Chinese models are usually the open weights ones so while there may be a training subsidy, the inference prices reflect real costs.
- wickedsight 2mo agoI get what you're saying, but I'm guessing that people asking how to pee is a drop in the bucket compared to the agentic loops being called to rename some variables across a project.
- runarberg 2mo agoI don’t use AI so this comes across to me as a bit of a culture shock, but is asking AI to rename a variable across project really something AI users are wasting their tokens on? In emacs I can do that with `S-l r r` or `M-x lsp-rename`. Using AI to do this seems extremely inefficient and wasteful, not to mention improper and unprofessional, and that is looking past the moral implication of training on stolen code and polluting our climate.