7 ms·
> Kimi K3 is Kimi’s most capable model to date, with 2.8 trillion parameters. This puts them on the top of the largest open models list: Kimi K3
by m3h 2mo ago
> Kimi K3 is Kimi’s most capable model to date, with 2.8 trillion parameters.
This puts them on the top of the largest open models list:
Kimi K3 2.8T
DeepSeek-V4-Pro 1.6T (49B active)
Kimi K2.6 ~1T (32B active)
GLM-5.2 754B (40B active)
DeepSeek-V3.2 685B
Mistral Large 3 675B
That's one mighty large model! Moonshot is going to need the USD 500 million reportedly raised earlier this year to run this model.
- wolttam 2mo agoI guess it remains to be seen whether this will be open-weights. We don't even know how many active params at this point.
- sudosysgen 2mo agoThe article says weights will be released in the coming days, and hints it's likely around 50-70B active params.
- wolttam 2mo agoIt did say that, but it doesn't any longer.
- simonw 2mo agoWhat's the URL of the article that used to say that?
- wolttam 2mo agohttps://platform.kimi.ai/docs/guide/kimi-k3-quickstart https://platform.kimi.ai/docs/guide/kimi-k3-quickstart this one, it used to have more information about the model itself, similar to the K2.6 and K2.7 pages. Edit: OpenRouter still describes it as an open-weight model: https://openrouter.ai/moonshotai/kimi-k3 https://openrouter.ai/moonshotai/kimi-k3 Guess we'll see!
- staticman2 2mo agoThat's a quickstart page for using the model on the platform not a page about the model. I am skeptical you are correct that it said something about model license earlier. Edited: I was wrong.
- InsideOutSanta 2mo agoNot the person you're responding to, just a person who still has the original version of the page open in their browser. Quoting from it: "Kimi K3 is the first open-source model to reach the 2.8-trillion-parameter scale. It is the latest step in Kimi's continued push of model-scale boundaries: in 9 of the past 12 months, Kimi models have set new records for open-source model scale." The page has definitely changed. (I'm not sure why you would be skeptical of somebody recollecting something they probably read only half an hour earlier.)
- staticman2 2mo agoI was skeptical because the 2.6 getting started description doesn’t say open source either. I do however appreciate the correction.
- all2 2mo agoIt would definitely be useful to save that off and upload it to archive.org.
- InsideOutSanta 2mo agoI saved the original text, but they've now reintroduced the exact same language in the new blog post.
- markasoftware 2mo agoRight now, if you search https://www.google.com/search?q=kimi+k3+open+weight https://www.google.com/search?q=kimi+k3+open+weight the blurb under the quickstart page contains the removed text.
- ignoramous 2mo agoThe full [Kimi K3] model weights will be released by July 27, 2026. Further details on the architecture, training, and evaluations will be released alongside the Kimi K3 technical report. https://archive.vn/KBzXr https://archive.vn/KBzXr
- anon373839 2mo agoThey are still describing it as open source: > Kimi K3 is the first open-source model to reach 2.8 trillion parameters. https://platform.kimi.ai/docs/guide/kimi-k3-quickstart https://platform.kimi.ai/docs/guide/kimi-k3-quickstart
- wolttam 2mo agoNice, they updated it again
- SwellJoe 2mo agoThe K3 marketing popup when I look at the Kimi Code page says "Kimi K3 Open Frontier Model". So, if it's not going to be open, they haven't told the whole team, yet.
- christophilus 2mo agoOpen like OpenAI, maybe.
- deleted 2mo ago[deleted]
- kroaton 2mo agoLing/Ring 1T-A50B and the new Inkling 975B-A41B deserve to be on that list.
- ignoramous 2mo ago> on the top of the largest open models list Moonshot (true to their name?) has always lead in terms of releasing the largest among open weight LLMs. > Moonshot is going to need the USD 500 million reportedly raised earlier this year to run this model. Think Moonshot, as a spin-out, can expect backing from its former parent, Alibaba? I don't think they would be particularly worried about finances, if the Kimi K series continues to outperform the Qwen Max series (which seems to be the case; while Kimi is also super popular in China).
- manquer 2mo ago> USD 500 million There was an another around after that . Moonshot raised $2B on a $20B valuation in May - https://techcrunch.com/2026/05/07/chinas-moonshot-ai-raises-2b-at-20b-valuation-as-demand-for-open-source-ai-skyrockets/ https://techcrunch.com/2026/05/07/chinas-moonshot-ai-raises-...
- toephu2 2mo agoHow many parameters do the top closed models use?
- yansongliu 2mo agoFable reportedly use 20T paramteres, 1T=1000B. Opus is probably 10T. That said, these are estimates based on model preformance and scope of general knowledge breadth.OpenAI, Anthropic, and Google have not openly report their model sizes. Chinese models are way behind on the mode size race due to lack of abudent AI infrustructures. That said, it seems Chinese models are going pretty well on a seprate route. They manage to achieve 80-90% performance with 1/10 of the model size. This is some what related to the diminishing reward situation described in the scaling law. I think it can also be attributed to their persistent research in this direction. Thinking and DSA (deepseek attention) were both developed and opensourced by Chinese labs then adopted worldwide.
- pixlmint 2mo agoCrazy how mighty GLM-5.2 is at less than a third the parameter count. Z.ai really cooked with that one.
- richjdsmith 2mo agoI was just looking at the same thing and how flawed it is that we tie parameter count with "intelligence". GLM-5.2 is my go to day to day model because of how darn good it is. I had no idea it had a substantially lower parameter count over deepseek v4.
- tom2026hn 2mo agoKimi has almost no advantage over Zhipu (Z.AI), so the performance boost likely comes from the number of parameters. The 2.8T model may not be as large as Fable, so Fable’s performance may also stem from the number of parameters. Or perhaps they quickly distilled Fable or Mythos. Distilling Mythos has a significant barrier to entry, and since Fable was released not long ago, is this even feasible? It outperforms Fable in several tests—how did Distilling achieve these results? Is this some kind of cross-vendor RSI (Recursive Self-Improvement) or RDI(Recursive Distillation Improvement)?
- oblio 2mo agoI wonder if we're a 5-10 years from running this on beefy consumer hardware, or more like 20+ years.