Head to head
Kimi K3 vs GLM-5.2
Which is better, Kimi K3 or GLM-5.2?
Kimi K3 finishes ahead of GLM-5.2 on the index, 87.5 to 87.1. It is the stronger of the two on Craft, while GLM-5.2 still leads on Speed, Control and Value. Both are scored on the same rubric, in the same week, and neither placing is sponsored.
Updated
Kimi K3
05 of 53
87.5
The largest open-weights model anyone has shipped, and it reads a million tokens without blinking.
GLM-5.2
06 of 53
87.1
Frontier-class output, an MIT licence, and a million tokens of context. That combination is the story.
The four scores, side by side
| Score | Kimi K3 | GLM-5.2 | Difference |
|---|---|---|---|
| Craft | 91 | 84 | +7 |
| Speed | 74 | 82 | +8 |
| Control | 91 | 95 | +4 |
| Value | 82 | 92 | +10 |
| Index score | 87.5 | 87.1 | +0.4 |
Which to pick
Pick Kimi K3 unless Speed, Control and Value is what decides it for you, which is exactly where GLM-5.2 is the better answer.
Kimi K3
Moonshot published K3 at a scale nobody else has open-sourced, with native vision, an explicit thinking mode and a million-token window. On long-document work and agentic tasks it is genuinely near the top of the field. Two caveats matter: the licence restricts commercial use above a revenue threshold, so it is open weights rather than open source, and the hosted API is priced like a frontier model rather than a Chinese one.
Where it shines
- Million-token context handled properly rather than nominally
- Native vision and an explicit thinking mode
- The strongest open-weights model on agentic benchmarks
GLM-5.2
GLM-5.2 is the model that makes the open-versus-closed argument concrete. It trades blows with the closed flagships on coding and general reasoning, it takes a million tokens of context, and the weights are published under MIT with no user threshold and no field-of-use clause. You can run it, tune it, ship it inside a product and never send a byte to anyone. Nothing else in this index offers that much capability on those terms.
Where it shines
- MIT licence with no user-count or field-of-use restrictions
- Competitive with closed flagships on coding benchmarks
- Million-token context on weights you can hold yourself