Kimi K2 scores 14.7 on the Respan Index (#122), ahead of gpt-oss-120b at 14.5 (#124). gpt-oss-120b leads on 2 of the 3 benchmarks both report. gpt-oss-120b has the larger context window (131K tokens).
| Benchmark | gpt-oss-120b | Kimi K2 |
|---|---|---|
| SimpleBench (SimpleBench) | 22.1% | 26.3% |
| WeirdML (Håvard Tveit Ihle) | 48.2% | 39.4% |
| AIME 2025 | 90% | 49.5% |
Higher is better. Each score links to where it was published.
* Estimated: no published scores in that category. How the Respan Index works