GPT-5.6 Luna scores 46.2 on the Respan Index (#33), ahead of Claude Sonnet 5 at 44.2 (#35). Claude Sonnet 5 leads on 15 of the 25 benchmarks both report. GPT-5.6 Luna is 8.9x cheaper per token at list price. GPT-5.6 Luna has the larger context window (1.1M tokens).
* Estimated: no published scores in that category. How the Respan Index works
| FrontierMath Tiers 1-3 v2 (Epoch AI) |
| 65.6% |
| 82.1% |
| GPQA Diamond | 90.5% | 92.3% |
| LiveBench Agentic Coding (LiveBench) | 59.4% | 48.4% |
| LiveBench Coding (LiveBench) | 80.7% | 82.9% |
| LiveBench Data Analysis (LiveBench) | 71.7% | 78% |
| LiveBench Instruction Following (LiveBench) | 63.9% | 60.1% |
| LiveBench Language (LiveBench) | 75% | 72.6% |
| LiveBench Mathematics (LiveBench) | 92.9% | 87.2% |
| LiveBench Reasoning (LiveBench) | 88.7% | 85.6% |
| Mystery Game Puzzles (Epoch AI) | 35% | 21% |
| OTIS Mock AIME 2024-2025 (Epoch AI) | 94.7% | 98.3% |
| SAGE (Vals AI) | 48.9% | 44.2% |
| SimpleQA Verified | 33.7% | 41% |
| SWE-rebench 2026-05-15 to 2026-07-01 (Nebius) | 56.8% | 43.6% |
| Terminal-Bench 4.0 (Vals AI) | 9.6% | 11.6% |
| Vending-Bench 2 (Andon Labs) | 6377.7 | 4094.71 |
| WeirdML (Håvard Tveit Ihle) | 68.8% | 60.9% |
| LMArena Agent (LMArena) | 0.044 | -0.0117 |
| LMArena Elo | 1462 | 1453 |
| LMArena Vision (LMArena) | 1265 | 1260 |
| SimpleBench (SimpleBench) | 60.6% | 46.8% |
Higher is better. Each score links to where it was published.