Llama 4 Maverick scores 13.6 on the Respan Index (#131), ahead of gpt-oss-20b at 11.7 (#133). gpt-oss-20b leads on 2 of the 2 benchmarks both report. Llama 4 Maverick has the larger context window (1M tokens).
Higher is better. Each score links to where it was published.
* Estimated: no published scores in that category. How the Respan Index works