BIG-Bench Hard
Twenty-three hard tasks from BIG-Bench. A 2022 benchmark, saturated for current models; kept for history.
As of 19 Sept 2026, Gemini 1.5 Pro 001 leads BIG-Bench Hard on BenchLeader with 89.2%, ahead of DeepSeek V3 at 87.5%, across 44 model configurations with a published result.
- Published by
- Epoch AI Benchmarking Hub
- Category
- Reasoning
- Index weight
- Reference only
- Models
- 44
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
Twenty-three hard tasks from BIG-Bench.
How it is scored
Accuracy, as compiled by Epoch AI from published results.
What to keep in mind
Saturated: frontier models score near the ceiling, so it separates only older or smaller models, and results are largely self-reported by developers.
- 1Gemini 1.5 Pro 00189.2%
- 2DeepSeek V387.5%
- 3Gemini 1.5 Pro 001 Feb2484.0%
- 4Llama 3.1 405B82.9%
- 5phi-3-medium 14B (medium)81.4%
- 6Qwen2.5 72B79.8%
- 7Phi 3 Small79.1%
- 8Deepseek v278.8%
- 9GPT 4 061375.1%
- 10Phi 3 Mini71.7%
- 11Yi-34B71.7%
- 12Stable Beluga 269.3%
- 13Llama 2-70B64.9%
- 14GPT-3.5 Turbo (older v0613)61.6%
- 15Phi-259.4%
| # | ||||
|---|---|---|---|---|
| 1 | 89.2% | 45.1 | 2024-05-14 | |
| 2 | 87.5% | 44.8 | 2024-12-26 | |
| 3 | 84.0% | – | 2024-02-15 | |
| 4 | 82.9% | 43.5 | 2024-07-23 | |
| 5 | 81.4% | – | 2024-04-23 | |
| 6 | 79.8% | 42.7 | 2024-09-19 | |
| 7 | 79.1% | – | 2024-04-23 | |
| 8 | 78.8% | – | 2024-05-07 | |
| 9 | 75.1% | 40.2 | 2023-06-13 | |
| 10 | 71.7% | 36.2 | 2024-04-23 | |
| 11 | 71.7% | 35.5 | 2023-11-22 | |
| 12 | 69.3% | – | 2023-07-20 | |
| 13 | 64.9% | 34.9 | 2023-07-18 | |
| 14 | 61.6% | – | 2023-06-13 | |
| 15 | 59.4% | – | 2023-12-12 | |
| 16 | 58.7% | – | 2024-02-26 | |
| 17 | 58.4% | – | 2023-02-24 | |
| 18 | 58.2% | 36.5 | 2023-07-18 | |
| 19 | 56.1% | 32.1 | 2023-09-27 | |
| 20 | 55.1% | – | 2024-02-21 | |
| 21 | 55.0% | – | 2023-09-24 | |
| 22 | 52.5% | – | 2023-09-18 | |
| 23 | 50.0% | – | 2023-02-24 | |
| 24 | 49.0% | – | 2023-09-06 | |
| 25 | 47.2% | – | 2023-09-06 | |
| 26 | 47.2% | – | 2023-11-22 | |
| 27 | 45.0% | – | 2023-09-28 | |
| 28 | 44.1% | – | 2023-07-18 | |
| 29 | 43.0% | – | 2023-04-12 | |
| 30 | 43.0% | – | 2023-07-11 | |
| 31 | 41.6% | – | 2023-09-20 | |
| 32 | 39.2% | 34.2 | 2023-07-18 | |
| 33 | 38.0% | – | 2023-06-22 | |
| 34 | 37.9% | – | 2023-02-24 | |
| 35 | 37.1% | – | 2023-03-15 | |
| 36 | 37.0% | – | 2023-07-05 | |
| 37 | 35.6% | – | 2023-05-05 | |
| 38 | 35.2% | – | 2024-02-21 | |
| 39 | 34.9% | – | 2024-11-29 | |
| 40 | 33.7% | – | 2023-06-24 | |
| 41 | 33.5% | – | 2023-02-24 | |
| 42 | 32.5% | – | 2023-06-01 | |
| 43 | 28.8% | – | 2023-04-24 | |
| 44 | 28.2% | – | 2023-11-30 |
Cite as: BenchLeader, “BIG-Bench Hard leaderboard”, https://www.benchleader.com/benchmarks/bbh, data as of 19 Sept 2026.
BIG-Bench Hard: questions
- What does BIG-Bench Hard measure?
- Twenty-three hard tasks from BIG-Bench. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads BIG-Bench Hard?
- Gemini 1.5 Pro 001 leads BIG-Bench Hard with 89.2% as of 19 Sept 2026, ahead of DeepSeek V3 at 87.5%.
- How many models have BIG-Bench Hard results?
- 44 model configurations have a BIG-Bench Hard result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
- Who runs BIG-Bench Hard and how often is it updated?
- BIG-Bench Hard is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does BIG-Bench Hard count toward the BenchLeader Index?
- No. BIG-Bench Hard is shown for reference but left out of the composite index.