BenchLeader

BIG-Bench Hard

Twenty-three hard tasks from BIG-Bench. A 2022 benchmark, saturated for current models; kept for history.

As of 19 Sept 2026, Gemini 1.5 Pro 001 leads BIG-Bench Hard on BenchLeader with 89.2%, ahead of DeepSeek V3 at 87.5%, across 44 model configurations with a published result.

Published by
Epoch AI Benchmarking Hub
Category
Reasoning
Index weight
Reference only
Models
44
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

Twenty-three hard tasks from BIG-Bench.

How it is scored

Accuracy, as compiled by Epoch AI from published results.

What to keep in mind

Saturated: frontier models score near the ceiling, so it separates only older or smaller models, and results are largely self-reported by developers.

44 of 44
#
1Gemini 1.5 Pro 001Google89.2%45.12024-05-14
2DeepSeek V3DeepSeekopen87.5%44.82024-12-26
3Gemini 1.5 Pro 001 Feb24Google84.0%2024-02-15
4Llama 3.1 405BMetaopen82.9%43.52024-07-23
5phi-3-medium 14BmediumMicrosoft81.4%2024-04-23
6Qwen2.5 72BAlibabaopen79.8%42.72024-09-19
7Phi 3 SmallMicrosoft79.1%2024-04-23
8Deepseek v2DeepSeekopen ↗78.8%2024-05-07
9GPT 4 0613OpenAI75.1%40.22023-06-13
10Phi 3 MiniMicrosoftopen71.7%36.22024-04-23
11Yi-34B01.AI71.7%35.52023-11-22
12Stable Beluga 2Stability AI69.3%2023-07-20
13Llama 2-70BMetaopen ↗64.9%34.92023-07-18
14GPT-3.5 Turbo (older v0613)OpenAI61.6%2023-06-13
15Phi-2Microsoft59.4%2023-12-12
16Nemotron-4 15BNVIDIA58.7%2024-02-26
17LLaMA-65BMetaopen58.4%2023-02-24
18Llama 2-13BMetaopen ↗58.2%36.52023-07-18
19Mistral 7B v0.3Mistral AIopen56.1%32.12023-09-27
20Gemma 7BGoogle55.1%2024-02-21
21Qwen-14BAlibaba55.0%2023-09-24
22internlm-20bShanghai AI Lab52.5%2023-09-18
23LLaMA-33BMeta50.0%2023-02-24
24Baichuan2-13BBaichuan49.0%2023-09-06
25Baichuan2-13BBaichuan47.2%2023-09-06
26Yi 6B01.AI47.2%2023-11-22
27Qwen-7BAlibaba45.0%2023-09-28
28Llama 2-34BMeta44.1%2023-07-18
29vicuna-13b-v1.1LMSYS43.0%2023-04-12
30Baichuan 1-13BBaichuan43.0%2023-07-11
31Baichuan 2-7BBaichuan41.6%2023-09-20
32Llama 2-7BMetaopen ↗39.2%34.22023-07-18
33MPT-30BMosaicML38.0%2023-06-22
34LLaMA-13BMeta37.9%2023-02-24
35Falcon-40BTII37.1%2023-03-15
36internlm-7bShanghai AI Lab37.0%2023-07-05
37MPT-7BMosaicML35.6%2023-05-05
38Gemma 2BGoogle35.2%2024-02-21
39INTELLECT-1Prime Intellect34.9%2024-11-29
40chatglm2-6bZhipu AI33.7%2023-06-24
41LLaMA-7BMeta33.5%2023-02-24
42Baichuan1-7BBaichuan32.5%2023-06-01
43Falcon-7BTII28.8%2023-04-24
44Qwen-1_8BAlibaba28.2%2023-11-30

Cite as: BenchLeader, “BIG-Bench Hard leaderboard”, https://www.benchleader.com/benchmarks/bbh, data as of 19 Sept 2026.

BIG-Bench Hard: questions

What does BIG-Bench Hard measure?
Twenty-three hard tasks from BIG-Bench. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads BIG-Bench Hard?
Gemini 1.5 Pro 001 leads BIG-Bench Hard with 89.2% as of 19 Sept 2026, ahead of DeepSeek V3 at 87.5%.
How many models have BIG-Bench Hard results?
44 model configurations have a BIG-Bench Hard result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
Who runs BIG-Bench Hard and how often is it updated?
BIG-Bench Hard is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
Does BIG-Bench Hard count toward the BenchLeader Index?
No. BIG-Bench Hard is shown for reference but left out of the composite index.