DeepSeek V4 Pro scores 46.9 on the Respan Index (#29), ahead of Muse Spark 1.2 at 46.6 (#31). DeepSeek V4 Pro leads on 11 of the 19 benchmarks both report. Both cost about the same per token at list price. Muse Spark 1.2 has the larger context window (1.0M tokens).
* Estimated: no published scores in that category. How the Respan Index works
| LiveBench Data Analysis (LiveBench) |
| 79.2% |
| 76.5% |
| LiveBench Instruction Following (LiveBench) | 67.7% | 74.3% |
| LiveBench Language (LiveBench) | 82.1% | 78.6% |
| LiveBench Mathematics (LiveBench) | 95.1% | 91.2% |
| LiveBench Reasoning (LiveBench) | 85.8% | 90% |
| SimpleQA Verified | 47% | 60.3% |
| Terminal-Bench 4.0 (Vals AI) | 14.1% | 6.1% |
| WeirdML (Håvard Tveit Ihle) | 66.2% | 60.3% |
| LMArena Agent (LMArena) | 0.0119 | -0.0327 |
| LMArena Elo | 1451 | 1494 |
| LMArena WebDev (LMArena) | 1583 | 1532 |
| DeepSWE v1.1 | 62.7% | 59.3% |
| Terminal-Bench 2.1 | 87.9% | 82.9% |
| Toolathlon-Verified (HKUST) | 74.4% | 75.9% |
| SimpleBench (SimpleBench) | 50.9% | 74.5% |
Higher is better. Each score links to where it was published.