Claude Fable 5.1 scores 69.4 on the Respan Index (#4), ahead of Grok 4.5 at 43.2 (#38). Claude Fable 5.1 leads on 23 of the 23 benchmarks both report. Grok 4.5 is 6.7x cheaper per token at list price. Claude Fable 5.1 has the larger context window (1M tokens).
* Estimated: no published scores in that category. How the Respan Index works
| Furniture Assembly (Epoch AI) |
| 70% |
| 22.5% |
| LiveBench Agentic Coding (LiveBench) | 66.1% | 56.5% |
| LiveBench Coding (LiveBench) | 86.4% | 68.6% |
| LiveBench Data Analysis (LiveBench) | 80.3% | 73% |
| LiveBench Instruction Following (LiveBench) | 73% | 71.5% |
| LiveBench Language (LiveBench) | 89.5% | 82.8% |
| LiveBench Mathematics (LiveBench) | 97% | 90.8% |
| LiveBench Reasoning (LiveBench) | 91.7% | 87.2% |
| OTIS Mock AIME 2024-2025 (Epoch AI) | 100% | 97.8% |
| SAGE (Vals AI) | 48.5% | 35% |
| SimpleQA Verified | 70.8% | 48.3% |
| Terminal-Bench 4.0 (Vals AI) | 58.1% | 8.6% |
| Vending-Bench 2 (Andon Labs) | 5421.56 | 3887.43 |
| WeirdML (Håvard Tveit Ihle) | 92.9% | 46.4% |
| LMArena Agent (LMArena) | 0.1431 | 0.0121 |
| LMArena Elo | 1501 | 1450 |
| LMArena Vision (LMArena) | 1288 | 1279 |
| LMArena WebDev (LMArena) | 1749 | 1552 |
| SimpleBench (SimpleBench) | 86.6% | 70% |
Higher is better. Each score links to where it was published.