Qwen3 235B-A22B scores 14.4 on the Respan Index (#125), ahead of GPT-4.1 mini at 14.3 (#127). GPT-4.1 mini leads on 3 of the 4 benchmarks both report. GPT-4.1 mini has the larger context window (1.0M tokens).
| Benchmark | GPT-4.1 mini | Qwen3 235B-A22B |
|---|---|---|
| GPQA Diamond | 65% | 70.7% |
| MATH Level 5 (Epoch AI) | 87.3% | 68.9% |
| WeirdML (Håvard Tveit Ihle) | 37.6% | 37.3% |
| LMArena Elo | 1383 | 1366 |
Higher is better. Each score links to where it was published.
* Estimated: no published scores in that category. How the Respan Index works