Llama 4 Maverick scores 13.6 on the Respan Index (#131), ahead of DeepSeek V3 at 12.4 (#132). Llama 4 Maverick leads on 3 of the 4 benchmarks both report. Llama 4 Maverick has the larger context window (1M tokens).
| Benchmark | DeepSeek V3 | Llama 4 Maverick |
|---|---|---|
| GPQA Diamond | 59.1% | 69.8% |
| SimpleBench (SimpleBench) | 18.9% | 27.7% |
| WeirdML (Håvard Tveit Ihle) | 36.1% | 24.5% |
| MMLU-Pro | 75.9% | 80.5% |
Higher is better. Each score links to where it was published.
* Estimated: no published scores in that category. How the Respan Index works