Claude Opus 5.5 scores 72.4 on the Respan Index (#1), ahead of Kimi K2.6 at 34.4 (#55). Claude Opus 5.5 leads on 17 of the 19 benchmarks both report. Kimi K2.6 is 4.7x cheaper per token at list price. Claude Opus 5.5 has the larger context window (1M tokens).
* Estimated: no published scores in that category. How the Respan Index works
| GPQA Diamond |
| 90.6% |
| 90.8% |
| LiveBench Agentic Coding (LiveBench) | 71.7% | 46.9% |
| LiveBench Coding (LiveBench) | 89.3% | 78.6% |
| LiveBench Data Analysis (LiveBench) | 80.3% | 65.1% |
| LiveBench Instruction Following (LiveBench) | 67% | 64.4% |
| LiveBench Language (LiveBench) | 86.3% | 75.1% |
| LiveBench Mathematics (LiveBench) | 97.1% | 84.3% |
| LiveBench Reasoning (LiveBench) | 92.2% | 79.4% |
| Mystery Game Puzzles (Epoch AI) | 71% | 18% |
| OTIS Mock AIME 2024-2025 (Epoch AI) | 100% | 96.1% |
| SAGE (Vals AI) | 45.8% | 50.2% |
| SimpleQA Verified | 72.2% | 34.9% |
| Vending-Bench 2 (Andon Labs) | 9235.25 | 6204.57 |
| LMArena Elo | 1504 | 1455 |
| LMArena WebDev (LMArena) | 1815 | 1509 |
Higher is better. Each score links to where it was published.