Claude Sonnet 4.6 and Kimi K2.6 tie on the Respan Index at 34.4. Claude Sonnet 4.6 leads on 11 of the 20 benchmarks both report. Kimi K2.6 is 3.5x cheaper per token at list price. Claude Sonnet 4.6 has the larger context window (1M tokens).
* Estimated: no published scores in that category. How the Respan Index works
| 77.9% |
| 65.1% |
| LiveBench Instruction Following (LiveBench) | 63.2% | 64.4% |
| LiveBench Language (LiveBench) | 76.1% | 75.1% |
| LiveBench Mathematics (LiveBench) | 87% | 84.3% |
| LiveBench Reasoning (LiveBench) | 84.8% | 79.4% |
| Mystery Game Puzzles (Epoch AI) | 16% | 18% |
| OTIS Mock AIME 2024-2025 (Epoch AI) | 85.8% | 96.1% |
| SAGE (Vals AI) | 46.6% | 50.2% |
| SimpleQA Verified | 35.5% | 34.9% |
| SWE-bench Verified (Epoch AI) | 75.2% | 76.7% |
| Vending-Bench 2 (Andon Labs) | 7204.14 | 6204.57 |
| WeirdML (Håvard Tveit Ihle) | 66.1% | 55.9% |
| LMArena Elo | 1458 | 1455 |
| LMArena Vision (LMArena) | 1275 | 1265 |
| LMArena WebDev (LMArena) | 1521 | 1509 |
| OSWorld-Verified (XLANG) | 72.1% | 73.1% |
Higher is better. Each score links to where it was published.