Muse Spark 1.3 scores 54.1 on the Respan Index (#13), ahead of Grok 4.6 at 51.0 (#19). Muse Spark 1.3 leads on 14 of the 20 benchmarks both report. Muse Spark 1.3 is 1.5x cheaper per token at list price. Muse Spark 1.3 has the larger context window (1.0M tokens).
* Estimated: no published scores in that category. How the Respan Index works
| LiveBench Agentic Coding (LiveBench) |
| 57% |
| 64.1% |
| LiveBench Coding (LiveBench) | 76.8% | 81.1% |
| LiveBench Data Analysis (LiveBench) | 73.9% | 79.6% |
| LiveBench Instruction Following (LiveBench) | 71.9% | 78% |
| LiveBench Language (LiveBench) | 83.7% | 82.8% |
| LiveBench Mathematics (LiveBench) | 92.6% | 95.9% |
| LiveBench Reasoning (LiveBench) | 90.5% | 89.7% |
| Mystery Game Puzzles (Epoch AI) | 34% | 25% |
| OTIS Mock AIME 2024-2025 (Epoch AI) | 99.2% | 99.2% |
| Terminal-Bench 4.0 (Vals AI) | 17.2% | 24.7% |
| LMArena Agent (LMArena) | 0.0128 | 0.04 |
| LMArena Elo | 1454 | 1494 |
| LMArena Vision (LMArena) | 1264 | 1290 |
| LMArena WebDev (LMArena) | 1620 | 1657 |
| SimpleBench (SimpleBench) | 75.9% | 81.8% |
| DeepSWE v1.1 | 65.9% | 75.4% |
Higher is better. Each score links to where it was published.