GLM-5.1 scores 34.4 on the Respan Index (#55), ahead of Kimi K2.7 Code at 34.1 (#58). They split the 10 benchmarks both report evenly. Kimi K2.7 Code is 1.3x cheaper per token at list price. Kimi K2.7 Code has the larger context window (262K tokens).
* Estimated: no published scores in that category. How the Respan Index works
| 93.3% |
| 95.6% |
| SimpleQA Verified | 34% | 36.5% |
| Vending-Bench 2 (Andon Labs) | 5634.41 | 5082.94 |
| WeirdML (Håvard Tveit Ihle) | 57.1% | 54.1% |
| LMArena WebDev (LMArena) | 1509 | 1473 |
| SimpleBench (SimpleBench) | 55.1% | 57.9% |
Higher is better. Each score links to where it was published.