BenchLeader

Benchmarks

87 benchmarks from 28 publishers. Each one feeds the category it belongs to; weights and exclusions are listed on the methodology page.

Reasoning

BenchmarkWhat it measuresModelsWeightSource
GPQA DiamondGraduate-level science questions written to be Google-proof. Run by Epoch AI.2721.0Epoch AI Benchmarking Hub
Humanity's Last Exam2,500 expert-written questions across 100+ subjects (Scale AI / CAIS).441.0Scale AI / CAIS
SimpleBenchTrick questions where humans score ~84%: spatio-temporal reasoning and social intelligence.851.0SimpleBench
LMArena Hard PromptsText arena rating on prompts judged hard.3500.5LMArena
LiveBench ReasoningLiveBench reasoning category.52LiveBench
GPQA Diamond (AA)Artificial Analysis' own GPQA Diamond run. Shown for reference; Epoch's run is in the index.525Artificial Analysis
Humanity's Last Exam (AA)Artificial Analysis' own HLE run. Shown for reference.524Artificial Analysis
GPQA Diamond (Vals)Vals AI's own GPQA Diamond run. Shown for reference.129Vals AI
EnigmaEvalPuzzle-hunt style multi-step reasoning. Scale AI.50.5Scale AI SEAL
Kagi LLM BenchmarkKagi's unpublished reasoning, coding and instruction questions.1350.5Kagi LLM Benchmark
ARC-AGI-1Semi-private ARC-AGI-1 evaluation set.1860.5ARC Prize
ARC-AGI-2Semi-private ARC-AGI-2 evaluation set.1871.0ARC Prize
ARC-AGI-3Semi-private ARC-AGI-3 interactive evaluation set.390.5ARC Prize

Coding

BenchmarkWhat it measuresModelsWeightSource
SWE-bench Verified (Epoch)Epoch AI's own SWE-bench Verified runs with a constant scaffold.330.5Epoch AI Benchmarking Hub
SciCodeResearch-grade scientific coding problems.1580.5SciCode
WeirdMLUnusual machine-learning tasks that require writing working training code.1500.5WeirdML
FrontierCodeHard real-world coding tasks run in agent harnesses (Cognition).290.5Cognition
GSO-BenchSoftware optimisation tasks measured against expert speed-ups.310.5GSO-Bench
LMArena CodingText arena rating on coding prompts.3450.5LMArena
LMArena WebDevPairwise votes on generated web apps.1181.0LMArena
LiveBench CodingLiveBench coding category.52LiveBench
AA Coding IndexArtificial Analysis' coding composite (Terminal-Bench Hard, SciCode). Shown for reference.0Artificial Analysis
SciCode (AA)Artificial Analysis' own SciCode run. Shown for reference.162Artificial Analysis
LiveCodeBenchCompetitive programming problems released after model cutoffs. Run by Vals AI.1351.0Vals AI
IOIInternational Olympiad in Informatics problems. Run by Vals AI.580.5Vals AI
SWE-bench (Vals)Vals AI's own SWE-bench run. Shown for reference.84Vals AI
SWE-Bench ProLong-horizon software engineering tasks in public repositories. Scale AI.231.0Scale AI SEAL
Aider PolyglotPercent of 225 Exercism problems solved on the second attempt.451.0Aider polyglot leaderboard
SWE-bench Verified (bash only)Official SWE-bench Verified results from the minimal bash-only mini-SWE-agent harness.441.0SWE-bench
SWE-bench Verified (any scaffold)Best official SWE-bench Verified result for the model with any agent scaffold. Shown for reference.68SWE-bench

Agents & tools

BenchmarkWhat it measuresModelsWeightSource
Terminal-BenchCommand-line agent tasks; best result per model across harnesses.661.0Terminal-Bench
OSWorld-Verified 2.0Computer-use tasks in a real desktop environment.140.5OSWorld
GDPvalWin rate against industry professionals on economically valuable tasks (OpenAI).110.5OpenAI
CybenchUnguided capture-the-flag cybersecurity tasks.210.5Cybench
Remote Labor IndexReal freelance projects completed end to end (Scale AI / CAIS).120.5Scale AI / CAIS
APEX-AgentsLong-horizon professional agent tasks (Mercor).630.5Mercor
LMArena AgentTask-outcome score (0–100) from real agentic sessions judged by their users.400.5LMArena
LiveBench Agentic CodingLiveBench agentic coding category.52LiveBench
AA Agentic IndexArtificial Analysis' agentic composite (τ²-Bench, Terminal-Bench Hard). Shown for reference.0Artificial Analysis
Terminal-Bench HardHard subset of Terminal-Bench in a fixed harness. Run by Artificial Analysis.3870.5Artificial Analysis
τ²-Bench Telecom (AA)Artificial Analysis' τ²-Bench telecom run. Shown for reference.392Artificial Analysis
Terminal-Bench 2.1 (Vals)Terminal-Bench 2.1 in Vals AI's fixed harness.610.5Vals AI
MCP AtlasReal-world tool use through the Model Context Protocol. Scale AI.301.0Scale AI SEAL
HiL-BenchWhether agents notice information gaps and ask clarifying questions. Scale AI.170.5Scale AI SEAL
BFCL OverallBerkeley Function Calling Leaderboard overall accuracy.831.0Berkeley Function Calling Leaderboard
τ²-benchMean pass^1 across the airline, retail and telecom domains.81.0τ²-bench

Maths

BenchmarkWhat it measuresModelsWeightSource
FrontierMath Tiers 1–3Unpublished research-level maths problems, tiers 1–3 (v2 private set). Run by Epoch AI.1031.0Epoch AI Benchmarking Hub
FrontierMath Tier 4The hardest FrontierMath tier: problems that take expert mathematicians days. Run by Epoch AI.590.5Epoch AI Benchmarking Hub
OTIS Mock AIMEMock AIME competition problems from 2024–2025. Run by Epoch AI.2550.5Epoch AI Benchmarking Hub
MATH Level 5Hardest tier of the MATH dataset. Largely saturated, kept at half weight. Run by Epoch AI.900.5Epoch AI Benchmarking Hub
ProofBenchFull written proofs graded for rigour (Vals AI).610.5Vals AI
LiveBench MathematicsLiveBench mathematics category.52LiveBench
AIME (Vals)Recent AIME problems, run by Vals AI at stated reasoning effort.910.5Vals AI
MGSMGrade-school maths in ten languages. Run by Vals AI.700.5Vals AI
AIME 2026AIME I and II 2026, evaluated on release by MathArena.311.0MathArena
HMMT February 2026Harvard-MIT Mathematics Tournament, February 2026. MathArena.311.0MathArena
IMO 2025International Mathematical Olympiad 2025, proof-graded. MathArena.60.5MathArena
MathArena ApexThe hardest problems from recent competitions. MathArena.460.5MathArena

Knowledge

BenchmarkWhat it measuresModelsWeightSource
SimpleQA VerifiedShort factual questions; measures recall without hallucination. Run by Epoch AI.751.0Epoch AI Benchmarking Hub
LiveBench Data AnalysisLiveBench data analysis category.52LiveBench
AA-OmniscienceKnowledge recall that penalises hallucination: correct answers minus confident wrong ones, from −100 to 100. Run by Artificial Analysis.4721.0Artificial Analysis
MMLU-ProHarder, reasoning-heavy successor to MMLU. Run by Vals AI.1291.0Vals AI
LegalBenchLegal reasoning tasks written by lawyers. Run by Vals AI.1360.5Vals AI
CorpFinQuestions over real corporate finance documents. Run by Vals AI.1250.5Vals AI
TaxEvalUS tax questions checked by professionals. Run by Vals AI.1370.5Vals AI
MedQAUS medical licensing exam questions. Run by Vals AI.900.5Vals AI
PRBench FinanceProfessional reasoning in finance, graded by practitioners. Scale AI.330.5Scale AI SEAL
PRBench LegalProfessional reasoning in legal practice, graded by practitioners. Scale AI.330.5Scale AI SEAL
MultiNRCNative multilingual reasoning across languages. Scale AI.420.5Scale AI SEAL

Instruction following

BenchmarkWhat it measuresModelsWeightSource
LiveBench LanguageLiveBench language category.52LiveBench
LiveBench Instruction FollowingLiveBench instruction-following category.0LiveBench
IFBenchPrecise instruction following on unseen constraints. Run by Artificial Analysis.4011.0Artificial Analysis
MultiChallengeMulti-turn instruction following under realistic conversation constraints. Scale AI.291.0Scale AI SEAL
TutorBenchTutoring tasks for high-school and AP subjects. Scale AI.270.5Scale AI SEAL

Human preference

BenchmarkWhat it measuresModelsWeightSource
LMArena TextOverall text arena rating with style control.3501.0LMArena
EQ-Bench 4Emotional intelligence in multi-turn role-play, judged pairwise into an Elo-style rating.270.5EQ-Bench

Multimodal

BenchmarkWhat it measuresModelsWeightSource
LMArena VisionPairwise votes on image-understanding prompts.1411.0LMArena
MMMU-ProCollege-level multimodal understanding. Run by Artificial Analysis.2461.0Artificial Analysis
VISTAVision-language understanding. Scale AI.571.0Scale AI SEAL
MMMU (validation)Official MMMU validation-set leaderboard.460.5MMMU
MMMU-Pro (official)Official MMMU-Pro leaderboard. Shown for reference.25MMMU

Long context

BenchmarkWhat it measuresModelsWeightSource
Fiction.LiveBench 120kLong-context comprehension of fiction at 120k tokens.351.0Fiction.live
AA-LCRLong-context reasoning across ~100k-token document sets. Run by Artificial Analysis.4661.0Artificial Analysis

Composite

BenchmarkWhat it measuresModelsWeightSource
Epoch Capabilities IndexEpoch AI's own composite capability score fitted across many benchmarks. Shown for reference; not part of the BenchLeader Index.235Epoch AI Benchmarking Hub
LiveBenchAverage of LiveBench's category scores on the current question set.521.0LiveBench
AA Intelligence IndexArtificial Analysis' composite of its own evaluations (GPQA, HLE, IFBench, LCR, τ²-Bench, Terminal-Bench Hard, SciCode, AA-Omniscience and more).5391.0Artificial Analysis
Vals IndexVals AI's composite across its benchmarks. Shown for reference.52Vals AI