Every model's score on every benchmark we track, each linked to where it was published. Independent runs and lab-reported numbers are kept apart. Many of these benchmarks feed the Respan Index.
177 of 177 benchmarks
Coding
3 models, higher is better
Datacurve's software engineering benchmark for coding agents. Datacurve also runs models itself on one harness and publishes the results as a live leaderboard.
Benchmark data from llmmetric.com, Respan's model data service.