gpt-oss-20b scores 11.7 on the Respan Index (#133), ahead of Qwen3 30B-A3B Instruct 2507 at 11.5 (#134). gpt-oss-20b leads on 3 of the 4 benchmarks both report. Qwen3 30B-A3B Instruct 2507 has the larger context window (262K tokens).
| Benchmark | gpt-oss-20b | Qwen3 30B-A3B Instruct 2507 |
|---|---|---|
| Chess Puzzles (Epoch AI) | 4% | 2% |
| GPQA Diamond | 71.5% | 55.6% |
| OTIS Mock AIME 2024-2025 (Epoch AI) | 65.3% | 62.2% |
| LMArena Elo | 1287 | 1383 |
Higher is better. Each score links to where it was published.
* Estimated: no published scores in that category. How the Respan Index works