Every model's score on every benchmark we track, each linked to where it was published. Independent runs and lab-reported numbers are kept apart. Many of these benchmarks feed the Respan Index.
177 of 177 benchmarks
Coding
28 models, higher is better
Datacurve's software engineering benchmark for coding agents. Datacurve also runs models itself on one harness and publishes the results as a live leaderboard.
Benchmark data from llmmetric.com, Respan's model data service.