DeepSeek V4 Pro scores 46.9 on the Respan Index (#29), ahead of Grok 4.5 at 43.2 (#38). Grok 4.5 leads on 14 of the 23 benchmarks both report. DeepSeek V4 Pro is 1.5x cheaper per token at list price. DeepSeek V4 Pro has the larger context window (1M tokens).
* Estimated: no published scores in that category. How the Respan Index works
| 45.3% |
| 57.2% |
| GPQA Diamond | 90.9% | 93.4% |
| LiveBench Agentic Coding (LiveBench) | 54.9% | 56.5% |
| LiveBench Coding (LiveBench) | 77.2% | 68.6% |
| LiveBench Data Analysis (LiveBench) | 79.2% | 73% |
| LiveBench Instruction Following (LiveBench) | 67.7% | 71.5% |
| LiveBench Language (LiveBench) | 82.1% | 82.8% |
| LiveBench Mathematics (LiveBench) | 95.1% | 90.8% |
| LiveBench Reasoning (LiveBench) | 85.8% | 87.2% |
| OTIS Mock AIME 2024-2025 (Epoch AI) | 96.7% | 97.8% |
| SimpleQA Verified | 47% | 48.3% |
| SWE-rebench 2026-05-15 to 2026-07-01 (Nebius) | 40.2% | 63.8% |
| Terminal-Bench 4.0 (Vals AI) | 14.1% | 8.6% |
| Vending-Bench 2 (Andon Labs) | 3284.52 | 3887.43 |
| WeirdML (Håvard Tveit Ihle) | 66.2% | 46.4% |
| LMArena Agent (LMArena) | 0.0119 | 0.0121 |
| LMArena Elo | 1451 | 1450 |
| LMArena WebDev (LMArena) | 1583 | 1552 |
| SimpleBench (SimpleBench) | 50.9% | 70% |
Higher is better. Each score links to where it was published.