gpt-oss-120b scores 14.5 on the Respan Index (#124), ahead of Claude 3.5 Sonnet at 14.4 (#125). gpt-oss-120b leads on 5 of the 6 benchmarks both report. Claude 3.5 Sonnet has the larger context window (200K tokens).
* Estimated: no published scores in that category. How the Respan Index works
| 3.3% |
| 90% |
| SWE-bench Verified | 49% | 62.4% |
Higher is better. Each score links to where it was published.