Benchmarks
87 benchmarks from 28 publishers. Each one feeds the category it belongs to; weights and exclusions are listed on the methodology page.
Reasoning
| Benchmark | What it measures | Models | Weight | Source |
|---|---|---|---|---|
| GPQA Diamond | Graduate-level science questions written to be Google-proof. Run by Epoch AI. | 272 | 1.0 | Epoch AI Benchmarking Hub |
| Humanity's Last Exam | 2,500 expert-written questions across 100+ subjects (Scale AI / CAIS). | 44 | 1.0 | Scale AI / CAIS |
| SimpleBench | Trick questions where humans score ~84%: spatio-temporal reasoning and social intelligence. | 85 | 1.0 | SimpleBench |
| LMArena Hard Prompts | Text arena rating on prompts judged hard. | 350 | 0.5 | LMArena |
| LiveBench Reasoning | LiveBench reasoning category. | 52 | – | LiveBench |
| GPQA Diamond (AA) | Artificial Analysis' own GPQA Diamond run. Shown for reference; Epoch's run is in the index. | 525 | – | Artificial Analysis |
| Humanity's Last Exam (AA) | Artificial Analysis' own HLE run. Shown for reference. | 524 | – | Artificial Analysis |
| GPQA Diamond (Vals) | Vals AI's own GPQA Diamond run. Shown for reference. | 129 | – | Vals AI |
| EnigmaEval | Puzzle-hunt style multi-step reasoning. Scale AI. | 5 | 0.5 | Scale AI SEAL |
| Kagi LLM Benchmark | Kagi's unpublished reasoning, coding and instruction questions. | 135 | 0.5 | Kagi LLM Benchmark |
| ARC-AGI-1 | Semi-private ARC-AGI-1 evaluation set. | 186 | 0.5 | ARC Prize |
| ARC-AGI-2 | Semi-private ARC-AGI-2 evaluation set. | 187 | 1.0 | ARC Prize |
| ARC-AGI-3 | Semi-private ARC-AGI-3 interactive evaluation set. | 39 | 0.5 | ARC Prize |
Coding
| Benchmark | What it measures | Models | Weight | Source |
|---|---|---|---|---|
| SWE-bench Verified (Epoch) | Epoch AI's own SWE-bench Verified runs with a constant scaffold. | 33 | 0.5 | Epoch AI Benchmarking Hub |
| SciCode | Research-grade scientific coding problems. | 158 | 0.5 | SciCode |
| WeirdML | Unusual machine-learning tasks that require writing working training code. | 150 | 0.5 | WeirdML |
| FrontierCode | Hard real-world coding tasks run in agent harnesses (Cognition). | 29 | 0.5 | Cognition |
| GSO-Bench | Software optimisation tasks measured against expert speed-ups. | 31 | 0.5 | GSO-Bench |
| LMArena Coding | Text arena rating on coding prompts. | 345 | 0.5 | LMArena |
| LMArena WebDev | Pairwise votes on generated web apps. | 118 | 1.0 | LMArena |
| LiveBench Coding | LiveBench coding category. | 52 | – | LiveBench |
| AA Coding Index | Artificial Analysis' coding composite (Terminal-Bench Hard, SciCode). Shown for reference. | 0 | – | Artificial Analysis |
| SciCode (AA) | Artificial Analysis' own SciCode run. Shown for reference. | 162 | – | Artificial Analysis |
| LiveCodeBench | Competitive programming problems released after model cutoffs. Run by Vals AI. | 135 | 1.0 | Vals AI |
| IOI | International Olympiad in Informatics problems. Run by Vals AI. | 58 | 0.5 | Vals AI |
| SWE-bench (Vals) | Vals AI's own SWE-bench run. Shown for reference. | 84 | – | Vals AI |
| SWE-Bench Pro | Long-horizon software engineering tasks in public repositories. Scale AI. | 23 | 1.0 | Scale AI SEAL |
| Aider Polyglot | Percent of 225 Exercism problems solved on the second attempt. | 45 | 1.0 | Aider polyglot leaderboard |
| SWE-bench Verified (bash only) | Official SWE-bench Verified results from the minimal bash-only mini-SWE-agent harness. | 44 | 1.0 | SWE-bench |
| SWE-bench Verified (any scaffold) | Best official SWE-bench Verified result for the model with any agent scaffold. Shown for reference. | 68 | – | SWE-bench |
Agents & tools
| Benchmark | What it measures | Models | Weight | Source |
|---|---|---|---|---|
| Terminal-Bench | Command-line agent tasks; best result per model across harnesses. | 66 | 1.0 | Terminal-Bench |
| OSWorld-Verified 2.0 | Computer-use tasks in a real desktop environment. | 14 | 0.5 | OSWorld |
| GDPval | Win rate against industry professionals on economically valuable tasks (OpenAI). | 11 | 0.5 | OpenAI |
| Cybench | Unguided capture-the-flag cybersecurity tasks. | 21 | 0.5 | Cybench |
| Remote Labor Index | Real freelance projects completed end to end (Scale AI / CAIS). | 12 | 0.5 | Scale AI / CAIS |
| APEX-Agents | Long-horizon professional agent tasks (Mercor). | 63 | 0.5 | Mercor |
| LMArena Agent | Task-outcome score (0–100) from real agentic sessions judged by their users. | 40 | 0.5 | LMArena |
| LiveBench Agentic Coding | LiveBench agentic coding category. | 52 | – | LiveBench |
| AA Agentic Index | Artificial Analysis' agentic composite (τ²-Bench, Terminal-Bench Hard). Shown for reference. | 0 | – | Artificial Analysis |
| Terminal-Bench Hard | Hard subset of Terminal-Bench in a fixed harness. Run by Artificial Analysis. | 387 | 0.5 | Artificial Analysis |
| τ²-Bench Telecom (AA) | Artificial Analysis' τ²-Bench telecom run. Shown for reference. | 392 | – | Artificial Analysis |
| Terminal-Bench 2.1 (Vals) | Terminal-Bench 2.1 in Vals AI's fixed harness. | 61 | 0.5 | Vals AI |
| MCP Atlas | Real-world tool use through the Model Context Protocol. Scale AI. | 30 | 1.0 | Scale AI SEAL |
| HiL-Bench | Whether agents notice information gaps and ask clarifying questions. Scale AI. | 17 | 0.5 | Scale AI SEAL |
| BFCL Overall | Berkeley Function Calling Leaderboard overall accuracy. | 83 | 1.0 | Berkeley Function Calling Leaderboard |
| τ²-bench | Mean pass^1 across the airline, retail and telecom domains. | 8 | 1.0 | τ²-bench |
Maths
| Benchmark | What it measures | Models | Weight | Source |
|---|---|---|---|---|
| FrontierMath Tiers 1–3 | Unpublished research-level maths problems, tiers 1–3 (v2 private set). Run by Epoch AI. | 103 | 1.0 | Epoch AI Benchmarking Hub |
| FrontierMath Tier 4 | The hardest FrontierMath tier: problems that take expert mathematicians days. Run by Epoch AI. | 59 | 0.5 | Epoch AI Benchmarking Hub |
| OTIS Mock AIME | Mock AIME competition problems from 2024–2025. Run by Epoch AI. | 255 | 0.5 | Epoch AI Benchmarking Hub |
| MATH Level 5 | Hardest tier of the MATH dataset. Largely saturated, kept at half weight. Run by Epoch AI. | 90 | 0.5 | Epoch AI Benchmarking Hub |
| ProofBench | Full written proofs graded for rigour (Vals AI). | 61 | 0.5 | Vals AI |
| LiveBench Mathematics | LiveBench mathematics category. | 52 | – | LiveBench |
| AIME (Vals) | Recent AIME problems, run by Vals AI at stated reasoning effort. | 91 | 0.5 | Vals AI |
| MGSM | Grade-school maths in ten languages. Run by Vals AI. | 70 | 0.5 | Vals AI |
| AIME 2026 | AIME I and II 2026, evaluated on release by MathArena. | 31 | 1.0 | MathArena |
| HMMT February 2026 | Harvard-MIT Mathematics Tournament, February 2026. MathArena. | 31 | 1.0 | MathArena |
| IMO 2025 | International Mathematical Olympiad 2025, proof-graded. MathArena. | 6 | 0.5 | MathArena |
| MathArena Apex | The hardest problems from recent competitions. MathArena. | 46 | 0.5 | MathArena |
Knowledge
| Benchmark | What it measures | Models | Weight | Source |
|---|---|---|---|---|
| SimpleQA Verified | Short factual questions; measures recall without hallucination. Run by Epoch AI. | 75 | 1.0 | Epoch AI Benchmarking Hub |
| LiveBench Data Analysis | LiveBench data analysis category. | 52 | – | LiveBench |
| AA-Omniscience | Knowledge recall that penalises hallucination: correct answers minus confident wrong ones, from −100 to 100. Run by Artificial Analysis. | 472 | 1.0 | Artificial Analysis |
| MMLU-Pro | Harder, reasoning-heavy successor to MMLU. Run by Vals AI. | 129 | 1.0 | Vals AI |
| LegalBench | Legal reasoning tasks written by lawyers. Run by Vals AI. | 136 | 0.5 | Vals AI |
| CorpFin | Questions over real corporate finance documents. Run by Vals AI. | 125 | 0.5 | Vals AI |
| TaxEval | US tax questions checked by professionals. Run by Vals AI. | 137 | 0.5 | Vals AI |
| MedQA | US medical licensing exam questions. Run by Vals AI. | 90 | 0.5 | Vals AI |
| PRBench Finance | Professional reasoning in finance, graded by practitioners. Scale AI. | 33 | 0.5 | Scale AI SEAL |
| PRBench Legal | Professional reasoning in legal practice, graded by practitioners. Scale AI. | 33 | 0.5 | Scale AI SEAL |
| MultiNRC | Native multilingual reasoning across languages. Scale AI. | 42 | 0.5 | Scale AI SEAL |
Instruction following
| Benchmark | What it measures | Models | Weight | Source |
|---|---|---|---|---|
| LiveBench Language | LiveBench language category. | 52 | – | LiveBench |
| LiveBench Instruction Following | LiveBench instruction-following category. | 0 | – | LiveBench |
| IFBench | Precise instruction following on unseen constraints. Run by Artificial Analysis. | 401 | 1.0 | Artificial Analysis |
| MultiChallenge | Multi-turn instruction following under realistic conversation constraints. Scale AI. | 29 | 1.0 | Scale AI SEAL |
| TutorBench | Tutoring tasks for high-school and AP subjects. Scale AI. | 27 | 0.5 | Scale AI SEAL |
Human preference
| Benchmark | What it measures | Models | Weight | Source |
|---|---|---|---|---|
| LMArena Text | Overall text arena rating with style control. | 350 | 1.0 | LMArena |
| EQ-Bench 4 | Emotional intelligence in multi-turn role-play, judged pairwise into an Elo-style rating. | 27 | 0.5 | EQ-Bench |
Multimodal
| Benchmark | What it measures | Models | Weight | Source |
|---|---|---|---|---|
| LMArena Vision | Pairwise votes on image-understanding prompts. | 141 | 1.0 | LMArena |
| MMMU-Pro | College-level multimodal understanding. Run by Artificial Analysis. | 246 | 1.0 | Artificial Analysis |
| VISTA | Vision-language understanding. Scale AI. | 57 | 1.0 | Scale AI SEAL |
| MMMU (validation) | Official MMMU validation-set leaderboard. | 46 | 0.5 | MMMU |
| MMMU-Pro (official) | Official MMMU-Pro leaderboard. Shown for reference. | 25 | – | MMMU |
Long context
| Benchmark | What it measures | Models | Weight | Source |
|---|---|---|---|---|
| Fiction.LiveBench 120k | Long-context comprehension of fiction at 120k tokens. | 35 | 1.0 | Fiction.live |
| AA-LCR | Long-context reasoning across ~100k-token document sets. Run by Artificial Analysis. | 466 | 1.0 | Artificial Analysis |
Composite
| Benchmark | What it measures | Models | Weight | Source |
|---|---|---|---|---|
| Epoch Capabilities Index | Epoch AI's own composite capability score fitted across many benchmarks. Shown for reference; not part of the BenchLeader Index. | 235 | – | Epoch AI Benchmarking Hub |
| LiveBench | Average of LiveBench's category scores on the current question set. | 52 | 1.0 | LiveBench |
| AA Intelligence Index | Artificial Analysis' composite of its own evaluations (GPQA, HLE, IFBench, LCR, τ²-Bench, Terminal-Bench Hard, SciCode, AA-Omniscience and more). | 539 | 1.0 | Artificial Analysis |
| Vals Index | Vals AI's composite across its benchmarks. Shown for reference. | 52 | – | Vals AI |