GPT-4.1 and Qwen3.5 35B-A3B tie on the Respan Index at 16.3. Qwen3.5 35B-A3B leads on 3 of the 4 benchmarks both report. Qwen3.5 35B-A3B is 5.1x cheaper per token at list price. GPT-4.1 has the larger context window (1.0M tokens).
| Benchmark | GPT-4.1 | Qwen3.5 35B-A3B |
|---|---|---|
| Chess Puzzles (Epoch AI) | 6% | 10% |
| GPQA Diamond | 66.3% | 83.5% |
| OTIS Mock AIME 2024-2025 (Epoch AI) | 38.3% | 54.4% |
| LMArena Elo | 1415 | 1396 |
Higher is better. Each score links to where it was published.
* Estimated: no published scores in that category. How the Respan Index works