Every model's score on every benchmark we track, each linked to where it was published. Independent runs and lab-reported numbers are kept apart. Many of these benchmarks feed the Respan Index.
177 of 177 benchmarks
Math
103 models, higher is better
45 competition-style problems written by students of the Olympiad Training for Individual Study (OTIS) program, with integer answers from 0 to 999. Epoch AI runs every model itself under one setup, so scores are directly comparable.
Benchmark data from llmmetric.com, Respan's model data service.