gpt-oss-20b scores 11.7 on the Respan Index (#133), ahead of GPT-4.1 nano at 9.4 (#135). gpt-oss-20b leads on 3 of the 4 benchmarks both report. GPT-4.1 nano has the larger context window (1.0M tokens).
| Benchmark | GPT-4.1 nano | gpt-oss-20b |
|---|---|---|
| GPQA Diamond | 50.3% | 71.5% |
| OTIS Mock AIME 2024-2025 (Epoch AI) | 28.9% | 65.3% |
| WeirdML (Håvard Tveit Ihle) | 19% | 40.9% |
| LMArena Elo | 1322 | 1287 |
Higher is better. Each score links to where it was published.
* Estimated: no published scores in that category. How the Respan Index works