GPT-4o scores 13.9 on the Respan Index (#130), ahead of Llama 4 Maverick at 13.6 (#131). They split the 2 benchmarks both report evenly. Llama 4 Maverick has the larger context window (1M tokens).
Higher is better. Each score links to where it was published.
* Estimated: no published scores in that category. How the Respan Index works