Every model's score on every benchmark we track, each linked to where it was published. Independent runs and lab-reported numbers are kept apart. Many of these benchmarks feed the Respan Index.
177 of 177 benchmarks
Coding
23 models, higher is better
Tests whether an AI agent can complete real tasks in a command-line terminal, such as building software, configuring systems and processing data, scored by automated checks.
Benchmark data from llmmetric.com, Respan's model data service.