Grok 4.5 scores 43.2 on the Respan Index (#38), ahead of Kimi K2.7 Code at 34.1 (#58). Grok 4.5 leads on 16 of the 19 benchmarks both report. Kimi K2.7 Code is 1.8x cheaper per token at list price. Grok 4.5 has the larger context window (500K tokens).
* Estimated: no published scores in that category. How the Respan Index works
| 57.2% |
| 54% |
| GPQA Diamond | 93.4% | 87.9% |
| LiveBench Agentic Coding (LiveBench) | 56.5% | 45.7% |
| LiveBench Coding (LiveBench) | 68.6% | 74% |
| LiveBench Data Analysis (LiveBench) | 73% | 62.7% |
| LiveBench Instruction Following (LiveBench) | 71.5% | 56.3% |
| LiveBench Language (LiveBench) | 82.8% | 77.9% |
| LiveBench Mathematics (LiveBench) | 90.8% | 79.6% |
| LiveBench Reasoning (LiveBench) | 87.2% | 82.8% |
| OTIS Mock AIME 2024-2025 (Epoch AI) | 97.8% | 95.6% |
| SimpleQA Verified | 48.3% | 36.5% |
| Vending-Bench 2 (Andon Labs) | 3887.43 | 5082.94 |
| WeirdML (Håvard Tveit Ihle) | 46.4% | 54.1% |
| LMArena WebDev (LMArena) | 1552 | 1473 |
| SimpleBench (SimpleBench) | 70% | 57.9% |
Higher is better. Each score links to where it was published.