Claude Opus 4.6 scores 42.1 on the Respan Index (#41), ahead of Gemini 3.1 Pro Preview at 41.4 (#42). Claude Opus 4.6 leads on 15 of the 28 benchmarks both report. Gemini 3.1 Pro Preview is 2.2x cheaper per token at list price. Gemini 3.1 Pro Preview has the larger context window (1.0M tokens).
* Estimated: no published scores in that category. How the Respan Index works
| FrontierMath Tiers 1-3 v2 (Epoch AI) |
| 66% |
| 59.6% |
| Furniture Assembly (Epoch AI) | 28.3% | 26.7% |
| GPQA Diamond | 91.3% | 94.3% |
| LiveBench Agentic Coding (LiveBench) | 49% | 44.1% |
| LiveBench Coding (LiveBench) | 78.2% | 76.5% |
| LiveBench Data Analysis (LiveBench) | 69.9% | 78.5% |
| LiveBench Instruction Following (LiveBench) | 63.3% | 79.1% |
| LiveBench Language (LiveBench) | 83.3% | 85.4% |
| LiveBench Mathematics (LiveBench) | 89.3% | 91% |
| LiveBench Reasoning (LiveBench) | 88.7% | 84% |
| Mystery Game Puzzles (Epoch AI) | 25% | 34% |
| OTIS Mock AIME 2024-2025 (Epoch AI) | 94.4% | 95.6% |
| SAGE (Vals AI) | 51.6% | 48.7% |
| SimpleQA Verified | 47% | 73.5% |
| SWE-bench Verified (Epoch AI) | 78.7% | 75.6% |
| Vending-Bench 2 (Andon Labs) | 8017.59 | 3774.25 |
| WeirdML (Håvard Tveit Ihle) | 78% | 72.1% |
| LMArena Elo | 1498 | 1480 |
| LMArena Vision (LMArena) | 1299 | 1279 |
| LMArena WebDev (LMArena) | 1546 | 1446 |
| AIME 2026 | 96.7% | 98.3% |
| tau2-bench Banking Knowledge (Sierra) | 27.3% | 26% |
| GSO Opt@1 (GSO) | 37.3% | 21.6% |
| SimpleBench (SimpleBench) | 67.6% | 79.6% |
Higher is better. Each score links to where it was published.