5 ms·
> In our evaluations, Kimi K3 delivers frontier-level performance. Among the models tested, its overall intelligence ranks second only to Claude Fable 5 and GPT
by ekojs 2mo ago
> In our evaluations, Kimi K3 delivers frontier-level performance. Among the models tested, its overall intelligence ranks second only to Claude Fable 5 and GPT-5.6 Sol. For the complete benchmark results, see our tech blog. The full model weights of Kimi K3 will be released in the coming days. More details on the architecture, training, and evaluation will be published together with the Kimi K3 technical report.
> K3 pushes the boundary of end-to-end knowledge work. On the GDPval-AA v2 leaderboard, Kimi K3 scores 1687. The benchmark evaluates AI models on real-world tasks across 44 occupations and 9 major industries; Kimi K3 ranks behind only Claude Fable 5 Max and GPT-5.6 Sol Max, and ahead of Claude Opus 4.8 Max at 1600.
> On AA-Briefcase, Kimi K3 scores 1527, ranking second among all models — behind only Claude Fable 5 Max and ahead of GPT-5.6 Sol Max (1495). AA-Briefcase is a private agentic knowledge-work benchmark developed by Artificial Analysis to evaluate frontier agentic capability in long-horizon knowledge work.
Really good benchmark score it seems. Maybe another DeepSeek moment right here.
- paxys 2mo ago> its overall intelligence ranks second only to Claude Fable 5 and GPT-5.6 Sol Pretty sure ranking “second” to two others means ranking third.
- ekojs 2mo agoYeah, bad wording it seems. Though a charitable interpretation is that Fable 5 and GPT 5.6 Sol are joint 1st place in the measurement.
- paxys 2mo agoDoesn’t matter, the next one is still third.
- cheesecakegood 2mo agoDENSE_RANK() vs RANK() claims another victim
- UqWBcuFx6NV4r 2mo agogirl, we get it, you can count. re-stating your point does not a conversation make.
- jnwatson 2mo agoIf there are two folks standing at gold, nobody gets the silver medal.
- worldthruword 2mo agoBut linearizing an equal magnitude quantities by alphabet priority would be unfair. Magnitude is the important quantity here.
- deleted 2mo ago[deleted]
- scotty79 2mo agoWhich is still great because it means neither of the two best financed labs in the world manage to produce even two models themselves that would beat Kimi K3.
- deleted 2mo ago[deleted]
- antonyt 2mo agoCharitably, you could read this as "its overall intelligence [is in a class that] ranks second only to [that of]..."
- ignoramous 2mo agoThis is actually what's meant but this bikeshed has been built for yak shaving.
- stingraycharles 2mo agoSince we’re bikeshedding: that’s not what yak shaving means.
- vl 2mo agoWhile you are technically correct, in English it’s perfectly fine to say it this way as well. “Second only” here has meaning “next after”, not “number two”.
- __mharrison__ 2mo agoSo... France took second to England and Argentina?
- vl 2mo agoFrance’s football team is second only to England’s and Argentina’s. It’s a miracle that in language same words have different meanings depending on context. If this wouldn’t be the case we could have hardcoded NLP algorithmically without inventing these expensive LLMs!
- make3 2mo agoSecond group essentially is how you have to think of it
- avazhi 2mo agoThat’s not what second means in this context in English, and it’s incorrect to use it that way. This is because for something to be second there must have been something in first and only first, and so on; in this case there was a first and a second already, and you cannot amalgamate then because they didn’t tie (and even if they did, they’d be 1 and 2). Both logically and grammatically, it’s incorrect.
- seba_dos1 2mo agoYou're both logically and grammatically wrong. You could even ask an LLM to explain the meaning of the phrase to you if you don't believe that.
- avazhi 2mo agoEither think and write for yourself or stay silent next time. It'd be infinitely better than telling another person to use an LLM to understand something you yourself don't understand and are too lazy to try to figure out. I wish you the best.
- deleted 2mo ago[deleted]
- krackers 2mo agoNot if the others tie for first place.
- Calazon 2mo agoStill third even then.
- Demiurge 2mo agoI think what's implied here, in colloquial terms, is that it's in the second tier.
- paxys 2mo agoExcept saying second tier would be bad marketing, so they decided to change the meaning of words instead.
- UqWBcuFx6NV4r 2mo agookay, so let’s get this straight: even though there are seemingly quite a few people here that clearly understand what is being said. However, the fact that YOU specifically either genuinely do not understand this / have never come across this before, or are being intentionally difficult because of some philosophical disagreement, feel that you can unilaterally assert that they’re “redefining words”? I don’t know if this is genuinely your first day on Earth or something, but if you’re trying to parse English like a programming language then you’re not only making things hard on yourself, but also 99% of people you’ll ever speak to.
- akoumjian 2mo agoWhere are you seeing this write up?
- ekojs 2mo agoI copied that from https://platform.kimi.ai/docs/guide/kimi-k3-quickstart https://platform.kimi.ai/docs/guide/kimi-k3-quickstart but it seems they updated the page to remove the benchmark score now.
- Aurornis 2mo ago> > K3 pushes the boundary of end-to-end knowledge work. On the GDPval-AA v2 leaderboard, Kimi K3 scores 1687. The benchmark evaluates AI models on real-world tasks across 44 occupations and 9 major industries; Kimi K3 ranks behind only Claude Fable 5 Max and GPT-5.6 Sol Max, and ahead of Claude Opus 4.8 Max at 1600. This is the same benchmark where Sonnet 5 outperforms Opus 4.8 max. Like all model releases, the benchmarks aren't going to tell the whole story. All of the open weight models come with amazing benchmark results now. It's hard to believe anything other than that the benchmarks are leaking into (or intentionally included) into training data.
- rd 2mo agoi’ll never really understand this comment. why would labs do this if they know private benchmark evals will come out in the next week?
- andai 2mo agoSonnet 5 does beat Opus 4.8 on several benchmarks. It just costs more and takes longer. (On several other benchmarks, it costs more, takes longer, and does worse.)
- ignoramous 2mo agoPossible, but pay-as-you-go Hy3 / DeepSeek v4 Pro / MiMo v2.5 Pro (from respective vendors) are genuinely good enough as daily drivers, given the costs (especially, low prices for input cache, which usually makes up 70%+ of total input for agentic workflows). I put in $10 in DeepSeek & Xiaomi MiMo, and I've barely used $1 each, in a week of coding work. Coding Plans by MiniMax ($20/mo for 1.7b tokens) and Z.ai (~$30/week use for $17/mo) are also tremendous value for money.
- simonw 2mo ago> In our evaluations, Kimi K3 delivers frontier-level performance What page does that come from? I'm having trouble tracking it down.
- wolttam 2mo agoIt was on the page linked in the top comment, but it's been removed.
- andai 2mo agoWhere is this from?
- deanc 2mo agoThat’s an interesting way to say you’re third. I’m only second to the ten other runners on my local Strava segments.
- adverbly 2mo ago> Maybe another DeepSeek moment right here. Surely not... What made DeepSeek disruptive was that the cost was 10X lower. In this case, the cost is about 2X lower the Sol I think? At 2X, you're pretty close to the error margins due to token efficiency etc... I'd say this is "on trend" for open models catching up to frontier labs, but its not a "change in the trend" like DeepSeek was IMO.
- hedora 2mo agoDeepSeek didn’t really change any trends though, unless you count the stock market. It was impressive work, but models were commoditizing and inference costs were dropping rapidly already. They were neither the first nor the last 10x optimization, from what I’ve seen.
- stavros 2mo agoIf you know of any other 10x optimisations currently, please let me know! I'm in the market for a model that's a tenth the price of a frontier model at the same level of quality.
- ac29 2mo agoYou do understand that the "frontier" people are usually talking about is the cost-intelligence frontier right? By definition there is no model that is both cheaper and as intelligent or better than another on the frontier.
- stavros 2mo agoOK, let me be more precise: If you know of a frontier model that's ten times cheaper than the previous frontier model at that level of intelligence, please let me know, I'm in the market for one. Is that better?
- ac29 2mo agoIts just a single benchmark, but Luna 5.6 xhigh scores within the margin of error the same as Opus 4.8 max on DeepSWE for 8x cheaper. Luna max is quite a bit higher than Opus and still 4x cheaper
- z3t4 2mo agoWhat hardware are they using to run the benchmarks?
- fastball 2mo agoIn my experience, the Chinese models are much more benchmaxxed than their frontier lab competitors, so I'm taking these results with a fairly large helping of salt.