6 ms·
Hopefully the recent work Moonshot did with Kimi K2.7 Code trickles in to the other open-model labs. Per AA, while K2.7 Code is roughly on par w/ K2.6 in terms
by h14h 3mo ago
Hopefully the recent work Moonshot did with Kimi K2.7 Code trickles in to the other open-model labs.
Per AA, while K2.7 Code is roughly on par w/ K2.6 in terms of intelligence, it uses half the output tokens to get there.
- h14h 3mo agoI've been doing some testing with GLM 5.2 on Fireworks and it looks like the "High" reasoning level uses fewer tokens than even K2.7 Code by a considerable margin (roughly half). Don't have any evals indicating how it compares on upper-bound quality, but for a well-defined task it seems like GLM 5.2 on "High" is remarkably token efficient. Looking forward to seeing where it lands on the AA index.