DeepSeek V3.1 scores 16.7 on the Respan Index (#115), ahead of GPT-4.1 at 16.3 (#117). DeepSeek V3.1 leads on 2 of the 3 benchmarks both report. GPT-4.1 has the larger context window (1.0M tokens).
| Benchmark | DeepSeek V3.1 | GPT-4.1 |
|---|---|---|
| SimpleBench (SimpleBench) | 40% | 27% |
| WeirdML (Håvard Tveit Ihle) | 38.4% | 39% |
| LMArena Elo | 1417 | 1415 |
Higher is better. Each score links to where it was published.
* Estimated: no published scores in that category. How the Respan Index works