Claude 3.7 Sonnet scores 17.0 on the Respan Index (#112), ahead of Qwen3 235B-A22B Thinking 2507 at 16.8 (#113). Qwen3 235B-A22B Thinking 2507 leads on 3 of the 3 benchmarks both report. Qwen3 235B-A22B Thinking 2507 is 8.0x cheaper per token at list price. Qwen3 235B-A22B Thinking 2507 has the larger context window (262K tokens).
| Benchmark | Claude 3.7 Sonnet | Qwen3 235B-A22B Thinking 2507 |
|---|---|---|
| GPQA Diamond | 79.7% | 80.1% |
| OTIS Mock AIME 2024-2025 (Epoch AI) | 57.8% | 86.7% |
| LMArena Elo | 1388 | 1400 |
Higher is better. Each score links to where it was published.
* Estimated: no published scores in that category. How the Respan Index works