Kimi K3 scores 50.6 on the Respan Index (#21), ahead of DeepSeek V4 Pro at 46.9 (#29). Kimi K3 leads on 22 of the 28 benchmarks both report. DeepSeek V4 Pro is 3.0x cheaper per token at list price. Kimi K3 has the larger context window (1.0M tokens).
* Estimated: no published scores in that category. How the Respan Index works
| 45.3% |
| 72.2% |
| GPQA Diamond | 90.9% | 93.5% |
| LiveBench Agentic Coding (LiveBench) | 54.9% | 62.2% |
| LiveBench Coding (LiveBench) | 77.2% | 81.4% |
| LiveBench Data Analysis (LiveBench) | 79.2% | 78.7% |
| LiveBench Instruction Following (LiveBench) | 67.7% | 71.4% |
| LiveBench Language (LiveBench) | 82.1% | 85.5% |
| LiveBench Mathematics (LiveBench) | 95.1% | 84.4% |
| LiveBench Reasoning (LiveBench) | 85.8% | 90.7% |
| Mystery Game Puzzles (Epoch AI) | 43% | 26% |
| OTIS Mock AIME 2024-2025 (Epoch AI) | 96.7% | 97.2% |
| SimpleQA Verified | 47% | 50.6% |
| Terminal-Bench 4.0 (Vals AI) | 14.1% | 17.2% |
| Vending-Bench 2 (Andon Labs) | 3284.52 | 5165.04 |
| WeirdML (Håvard Tveit Ihle) | 66.2% | 82.6% |
| LMArena Agent (LMArena) | 0.0119 | 0.0418 |
| LMArena Elo | 1451 | 1488 |
| LMArena WebDev (LMArena) | 1583 | 1658 |
| AIME 2026 | 96.7% | 96.7% |
| DeepSWE v1.1 | 62.7% | 67.3% |
| Humanity's Last Exam (with tools) | 60% | 56% |
| Terminal-Bench 2.1 | 87.9% | 88.3% |
| Toolathlon-Verified (HKUST) | 74.4% | 76.5% |
| SimpleBench (SimpleBench) | 50.9% | 60.7% |
Higher is better. Each score links to where it was published.