Every model's score on every benchmark we track, each linked to where it was published. Independent runs and lab-reported numbers are kept apart. Many of these benchmarks feed the Respan Index.
177 of 177 benchmarks
Math
57 models, higher is better
Original, unpublished mathematics problems created and verified by expert mathematicians: 295 base problems (Tiers 1-3) plus 43 exceptionally hard Tier 4 problems. Answers are checked automatically, and Epoch AI evaluates models on the private set.
Benchmark data from llmmetric.com, Respan's model data service.