Qwen3 235B-A22B Thinking 2507 scores 16.8 on the Respan Index (#113), ahead of DeepSeek V3.1 at 16.7 (#115). They split the 2 benchmarks both report evenly. Qwen3 235B-A22B Thinking 2507 has the larger context window (262K tokens).
| Benchmark | DeepSeek V3.1 | Qwen3 235B-A22B Thinking 2507 |
|---|---|---|
| WeirdML (Håvard Tveit Ihle) | 38.4% | 41% |
| LMArena Elo | 1417 | 1400 |
Higher is better. Each score links to where it was published.
* Estimated: no published scores in that category. How the Respan Index works