7 ms·
The numbers from Anthropic seem heavily cherry-picked, Artificial Analysis has Opus 5 at 1.25x the cost of Sonnet and 2x the cost of GPT 5.6 and K3. https://ar
by x313 2mo ago
The numbers from Anthropic seem heavily cherry-picked, Artificial Analysis has Opus 5 at 1.25x the cost of Sonnet and 2x the cost of GPT 5.6 and K3.
https://artificialanalysis.ai/?cost=cost-per-task https://artificialanalysis.ai/?cost=cost-per-task
- reinitctxoffset 2mo ago[flagged]
- idiotsecant 2mo agoI feel like I am having a stroke. What is this
- reinitctxoffset 2mo ago[flagged]
- throwawayoaky 2mo agolooks like the agent-judged results of an agent-built 'eval' based on some examples derived from this person's real work. and clearly part of a larger document. this kind of slop is kind of useful but opus 4 was the first generation that was any good at writing its own prompts/evals/rubrics so there's a certain sloop to it..
- SwellJoe 2mo agoI don't understand how the K3 numbers keep coming out cheap for people. I recently started to add it to my security auditing benchmarks and found it was going to cost about twice as much as Opus 4.8. It blew through the $100 budget I'd set at like 11%. In the tasks I'm doing it seems crazy expensive because it chews so much, burning a tremendous amount of tokens.
- InsideOutSanta 2mo agoI think the way people usually compare pricing is fundamentally flawed. You can't compare token prices because different models use different tokenizers, and you can't compare tokenizer-normalized token prices because different models at different settings use more or fewer tokens to complete the same task at a different level of quality. Based on my entirely subjective experience, the $100 Moonshot plan using only K3 is comparable to the $200 Anthropic deal using the whole Fable allocation and Opus 4.8 for the rest.
- SwellJoe 2mo agoI got the $19 plan, and it's anemic. One tiny task blew through the 5-hour budget and 19% of the weekly budget. A completely useless amount of usage. OpenAI's $20 plan feels like 100x more generous (I don't think I'm exaggerating here). Someone in another thread said their plans are cheaper in China, maybe that's the difference, I dunno. But, I'm finding Kimi K3 terrifyingly expensive in the way that Fable and GPT 5.5 Pro are at token rates. Not as expensive as those, but expensive enough to where if you don't put a budget cap on it, you might wake up bankrupt if you leave a task running overnight. Not because of the per-token cost, but because how many tokens it's going to burn.
- try-working 2mo agoI have the second largest Kimi plan, the Chinese version. When K2.6 was their latest model, the quota was good; it was like GPT $100 is now or what the $20 version was in December. When K2.7 was released, they cut quota by 80%. I can't tell how much they have further cut it after the K3 release because it's barely worth using at all. I just use it in my model router since I have the annual plan paid for. It's just not a serious model or company.
- InsideOutSanta 2mo agoYes, the OpenAI plans are much more generous than both Moonshot's and Anthropic's. It's the only provider of the three where the $20 plan is at all usable for programming.
- adgjlsfhk1 2mo agoTesting at max effort likely doesn't produce optimal results.
- sggyamg 2mo agoCan you be more explicit?
- adgjlsfhk1 2mo agoMax effort is the way to give the highest perf, but not highest perf/$. Having claude (or other models) use a lower effort can often be 80% as smart but get to the results 10x faster for the problems where it works.
- crimist 2mo agoWe've seen models perform worse at higher efforts in our vuln detection evals. For example IIRC gpt 5.5 and 5.6 both scored better or high as compared to xhigh.
- deleted 2mo ago[deleted]