DeepSeek V4 Pro scores 46.9 on the Respan Index (#29), ahead of Kimi K2.6 at 34.4 (#55). DeepSeek V4 Pro leads on 16 of the 21 benchmarks both report. Kimi K2.6 is 1.2x cheaper per token at list price. DeepSeek V4 Pro has the larger context window (1M tokens).
* Estimated: no published scores in that category. How the Respan Index works
| 54.9% |
| 46.9% |
| LiveBench Coding (LiveBench) | 77.2% | 78.6% |
| LiveBench Data Analysis (LiveBench) | 79.2% | 65.1% |
| LiveBench Instruction Following (LiveBench) | 67.7% | 64.4% |
| LiveBench Language (LiveBench) | 82.1% | 75.1% |
| LiveBench Mathematics (LiveBench) | 95.1% | 84.3% |
| LiveBench Reasoning (LiveBench) | 85.8% | 79.4% |
| Mystery Game Puzzles (Epoch AI) | 43% | 18% |
| OTIS Mock AIME 2024-2025 (Epoch AI) | 96.7% | 96.1% |
| SimpleQA Verified | 47% | 34.9% |
| SWE-bench Verified (Epoch AI) | 77.6% | 76.7% |
| Vending-Bench 2 (Andon Labs) | 3284.52 | 6204.57 |
| WeirdML (Håvard Tveit Ihle) | 66.2% | 55.9% |
| LMArena Elo | 1451 | 1455 |
| LMArena WebDev (LMArena) | 1583 | 1509 |
| AIME 2026 | 96.7% | 95.8% |
| Toolathlon-Verified (HKUST) | 74.4% | 58% |
Higher is better. Each score links to where it was published.