LAMBADA
Predicting the last word of a passage. A 2016 benchmark, saturated for current models; kept for history.
As of 19 Sept 2026, Megatron-Turing NLG 530B leads LAMBADA on BenchLeader with 87.2%, ahead of InstructGPT 175B at 86.4%, across 24 model configurations with a published result.
- Published by
- Epoch AI Benchmarking Hub
- Category
- Knowledge
- Index weight
- Reference only
- Models
- 24
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
Predicting the last word of a passage.
How it is scored
Accuracy, as compiled by Epoch AI from published results.
What to keep in mind
Saturated: frontier models score near the ceiling, so it separates only older or smaller models, and results are largely self-reported by developers.
- 1Megatron-Turing NLG 530B87.2%
- 2InstructGPT 175B86.4%
- 3Falcon-180B79.8%
- 4Llama 2-70B78.9%
- 5Inflection-178.5%
- 6Palm 540B77.9%
- 7LLaMA-65B77.7%
- 8Falcon-40B77.3%
- 9LLaMA-33B77.2%
- 10Llama 2-13B76.5%
- 11LLaMA-13B75.2%
- 12Falcon-7B74.9%
- 13Baichuan2-13B74.0%
- 14Llama 2-7B73.3%
- 15Baichuan 2-7B73.3%
| # | ||||
|---|---|---|---|---|
| 1 | 87.2% | – | 2021-10-11 | |
| 2 | 86.4% | – | 2022-01-27 | |
| 3 | 79.8% | – | 2023-09-06 | |
| 4 | 78.9% | 34.9 | 2023-07-18 | |
| 5 | Inflection-1Inflection AI | 78.5% | – | 2023-06-22 |
| 6 | 77.9% | – | 2022-04-04 | |
| 7 | 77.7% | – | 2023-02-24 | |
| 8 | 77.3% | – | 2023-03-15 | |
| 9 | 77.2% | – | 2023-02-24 | |
| 10 | 76.5% | 36.5 | 2023-07-18 | |
| 11 | 75.2% | – | 2023-02-24 | |
| 12 | 74.9% | – | 2023-04-24 | |
| 13 | 74.0% | – | 2023-09-06 | |
| 14 | 73.3% | 34.2 | 2023-07-18 | |
| 15 | 73.3% | – | 2023-09-20 | |
| 16 | 73.3% | – | 2023-02-24 | |
| 17 | 71.8% | – | 2023-09-18 | |
| 18 | 71.3% | – | 2023-07-20 | |
| 19 | 71.1% | – | 2023-09-28 | |
| 20 | 70.0% | – | 2023-05-05 | |
| 21 | 67.9% | – | 2023-09-28 | |
| 22 | 67.0% | – | 2023-07-05 | |
| 23 | 58.4% | – | 2023-11-30 | |
| 24 | 54.3% | – | 2023-06-24 |
Cite as: BenchLeader, “LAMBADA leaderboard”, https://www.benchleader.com/benchmarks/lambada, data as of 19 Sept 2026.
LAMBADA: questions
- What does LAMBADA measure?
- Predicting the last word of a passage. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads LAMBADA?
- Megatron-Turing NLG 530B leads LAMBADA with 87.2% as of 19 Sept 2026, ahead of InstructGPT 175B at 86.4%.
- How many models have LAMBADA results?
- 24 model configurations have a LAMBADA result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
- Who runs LAMBADA and how often is it updated?
- LAMBADA is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does LAMBADA count toward the BenchLeader Index?
- No. LAMBADA is shown for reference but left out of the composite index.