6 ms·
Some official benchmark numbers posted in Chinese social media (I am sure they will publish an English blogpost later too): https://mp.weixin.qq.com/s/V4xhEIy8
by natrys 2mo ago
Some official benchmark numbers posted in Chinese social media (I am sure they will publish an English blogpost later too):
https://mp.weixin.qq.com/s/V4xhEIy8xDXSMDPrPkmUAQ https://mp.weixin.qq.com/s/V4xhEIy8xDXSMDPrPkmUAQ
Generally looks like a Sol/Fable tier model, better across the board than Opus 4.8.
(Edit) English blogpost is up now: https://www.kimi.com/blog/kimi-k3 https://www.kimi.com/blog/kimi-k3
- EugeneOZ 2mo agoAny benchmark where Sol is better than Fable at coding is ridiculous.
- zarzavat 2mo agoIt's like reading Anthropic's obituary.
- refulgentis 2mo agoFable is by Anthropic, and this is too expensive, GLM 5.2 is roughly the same quality at a much cheaper price. (I mantain a client with llama.cpp and 101 models across 14 companies by http)
- LaurensBER 2mo agoAs much as I like GLM 5.2 it's clearly a step below Opus (or even Fable) for more complicated tasks. I would place it at Opus 4.6/4.7 level. Having said that, the safety system on Fable makes it an extremely unattractive model. It feels that half of the time you're paying double for Opus level performance.
- jml78 2mo agoFable won’t even generate a jwt to test endpoints because it is security related. It is crazy capable but useless for real work
- weego 2mo agoUnless your real work is outside the scope of one tiny niche of work.
- arcanemachiner 2mo agoEh, it doesn't hit you until it hits you. I finally bumped into a task that Codex would refuse to work on. Was I attempting to reverse-engineer a GPU driver? Yes. Was I trying to hack into the DoD? No. I wasn't doing anything wrong, but that's not what OpenAI's safety mechanisms thought.
- pimeys 2mo agoGLM has issues with tool calls and nested JSON and it wastes tokens pretty often. I see it being a bit above half the price of Opus in a bit more complex eval tasks. With some RL you could probably get the tool calls sorted and the price down.
- austinthetaco 2mo agoThis is weird and reactionary. Lots of organizations are continuing to refuse to use chinese models due to security and IP concerns. Anthropic/american models aren't going anywhere anytime soon.
- sscaryterry 2mo agoNope, but I think this is maybe the critical mass needed to finally crash the AI hype/datacenter cost problem everyones is talking about. With Oracle being junk before this, more will follow.
- ai-x 2mo agoModels need datacenters to run. It also need other services to do anything useful
- sscaryterry 2mo agoThe point: Fable isn't worth what Anthropic says it is, so Anthropic isn't as valuable as they make themselves out to be. The DeepSeek incident has already shown it, this is a reminder.
- stevefan1999 2mo agoOracle is fine, it's just that they can't really expect political decisions that hindered it to accquire TikTok which will be slated to be the biggest customer if the deal went through. Now they are betting with Project Stargate but it also seems to be crumbling down. But don't forget that they literally hold the biggest databases, both in commercial and open source, that is, Oracle Database and MySQL. Plus Oracle Java they literally controls at least 30% of the internet's software infrastructure. And also with a good team of attorneies enforcing the licenses, they can squeeze so much money at the cost of morality. Also recently they downgraded the always free OCI ARM instance from 4C24G to 2C12G without telling anyone.
- re-thc 2mo ago
- scrollop 2mo agoNah: https://www.youtube.com/watch?v=LSlV206xPqM https://www.youtube.com/watch?v=LSlV206xPqM These real world examples show it's one tier away.
- deleted 2mo ago[deleted]
- VulgarExigency 2mo agoThese "real world" examples are nothing like the way I use LLMs from within a harness. GPT 5.6 Sol and Fable are clearly more impressive, but how does this translate to interactive agent use, or use under an agent orchestration framework?
- pimeys 2mo agoThis is a question I am going to get an answer tomorrow with evals. Extremely interesting...
- zarzavat 2mo agoIf Chinese AI companies can train a model that's slightly worse than the frontier, then there's no reason why they can't train a model that is slightly better than the frontier. Everybody can agree that K3 doesn't clearly surpass Fable. However, inevitably there will be a time in the future when a Chinese AI company releases a model that's better than any US model. K3 isn't the knockout blow but it's the 2nd knockdown that makes everyone in the arena realize that the fighter is not winning the fight.
- threatripper 2mo agoAnthropic is arguably still better in tooling and integrating model and tooling. Good habit beats raw intelligence. For code editing Cursor editor tooling is even better.
- baq 2mo agoCertainly for their IPO, anyway
- GodelNumbering 2mo agoThe link has 6 well-known benchmarks where this beats Fable (out of 14 I counted). If the numbers hold up scrutiny, this is scary good. Forget about their pricing but the companies that do have means to host such models fully on-prem are also the same companies that are paying tens of millions of $ in inference cost every month, and are by extension the biggest customers of OAI and Anthropic
- echelon 2mo agoOpen Source >>> Closed Source [1] I don't want to cheer against my country, but we've given up on open source. The way Anthropic and OpenAI treat their customers as adversaries is embarrassing. I will cheer for China, for Kimi, and for z.ai until we have something in the same category. [1] I'd even be fine with open weights, fair source, or anything that let us have direct access to the weights. Even if that came with stipulations. Don't hide the weights from us.
- GodelNumbering 2mo agoI am with you in the spirit of openweights but I am trying to hard-avoid bringing countries into this. The narrative of US vs China only benefits those who want regulatory capture in the US since attacking China is politically much easier than attacking open-weights, so certain groups like to repeatedly call them 'Chinese models'.
- echelon 2mo agoIt's much more a rallying cry for open weights funding than it is for regulatory capture. The argument on our side wins - if America or the West don't do open source, China will. And that means -- with certainty -- that China wins the market. Every politician and VC should hear that loud and clear.
- titanomachy 2mo agoI call them “Chinese models” without vitriol. I think they’re great. Although it’s good to remember that all models have ideological biases inherited from their creators.
- tgtweak 2mo agoI think given how much benchmaxxing we're seeing - the anecdotal evidence of how competent this model is (and efficient) will depend on user's actual real-world use cases. Given the pricing, it suggests that this model is much more efficient/competent than previous-gen OS/distilled models.
- vonneumannstan 2mo ago[flagged]
- MaxPock 2mo agoDo you have moat if your advanced model can be distilled in a month or two ?
- twobitshifter 2mo agodistillation attack? why the violent word choice? When OpenAI crawled Github was that an attack?
- Fizz43 2mo agoyes
- rowanG077 2mo agoDistillation is not an attack. It simply a way to train a model. Not doing it when you are behind is akin to snatching defeat from the jaws of victory.
- mensetmanusman 2mo agoIt is an attack at a sufficient level of sophisticated analysis. If you destroy the game theoretic first mover advantage, then you destroy the economic incentive to improve things.
- rowanG077 2mo agoGiven that model distillation has existed since the early days of the current AI boom, and no robust defense has been demonstrated, the available evidence does not support your theory.
- ptnpzwqd 2mo agoBe that as it may, it would seem absurd if we start calling distillation out as antagonistic, but don't do the same for the SOTA models being trained on human-created data.
- scrollop 2mo agoMeh, not fable/sol tier: https://www.youtube.com/watch?v=LSlV206xPqM https://www.youtube.com/watch?v=LSlV206xPqM
- natrys 2mo agoIf anecdote is data, then here's another point: https://nitter.net/synthwavedd/status/2077537805715005724#m https://nitter.net/synthwavedd/status/2077537805715005724#m (As an aside, I don't know how it was professional of Arena to unmask an unreleased cloaked model on their platform. Also practically, upstream could have been A/B testing multiple variants under same endpoint, casting validity of such pre-announcement tests into question)
- EugeneOZ 2mo agoWhy we should waste 46 minutes instead of briefly looking at charts for 20 seconds? To pay their ads? No, thanks.