6 ms·
(2025) As it's 9 months old and they just had a major model release
by jasonjmcghee 2mo ago
(2025)
As it's 9 months old and they just had a major model release
- GaggiX 2mo agoI believe OP posted it because the new Kimi K3 has 69 KDA layers (the rest are 24 Gated MLA), I think previous large Kimi models had only MLA layers.
- yorwba 2mo agoIt's not the same KDA as used in Kimi Linear, though.
- throwa356262 2mo agoFor K3 read this instead: https://arxiv.org/abs/2607.24653 https://arxiv.org/abs/2607.24653 The main contribution of the K3 paper is Stable LatentMoE. Like some other models it compresses data sent between layers, which puts certain requirements on the router. K3 improves performance by using a more balanced expert selection strategy.
- mcbuilder 2mo agoCompared to the Opus 5 "model card", which read like a standard Anthropic set of alignment principles and safety concerns, this presents a plethora of useful technical details that advances the state of the art.
- throwa356262 2mo agoSame with DeepSeek papers, they are a joy to read.
- senko 2mo agoNot an expert, but looks like they did a lot more work on the RL part (9 expert models, full sandbox access for agentic tasks, etc)?
- verdverm 2mo agomost new effort in training comes in the late phase with RL techniques the pretraining (slurping the internet) only goes so far, the new data being used is from human preferences and agent traces (designed and/or distilled)
- cptcobalt 2mo agoRather under-discussed back then: https://news.ycombinator.com/item?id=45766937 https://news.ycombinator.com/item?id=45766937