GPT-4o mini scores 8.8 on the Respan Index (#136), ahead of Llama 3.3 70B Instruct at 8.6 (#137). Llama 3.3 70B Instruct leads on 4 of the 7 benchmarks both report. Llama 3.3 70B Instruct has the larger context window (131K tokens).
* Estimated: no published scores in that category. How the Respan Index works
| 11.8% |
| 14.4% |
| LMArena Elo | 1318 | 1274 |
| BALROG (BALROG) | 17.4% | 23% |
Higher is better. Each score links to where it was published.