TriviaQA
Trivia questions with evidence documents. A 2017 benchmark, saturated for current models; kept for history.
As of 19 Sept 2026, Llama 2-70B leads TriviaQA on BenchLeader with 87.6%, ahead of Claude 2 at 87.5%, across 36 model configurations with a published result.
- Published by
- Epoch AI Benchmarking Hub
- Category
- Knowledge
- Index weight
- Reference only
- Models
- 36
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
Trivia questions with evidence documents.
How it is scored
Accuracy, as compiled by Epoch AI from published results.
What to keep in mind
Saturated: frontier models score near the ceiling, so it separates only older or smaller models, and results are largely self-reported by developers.
- 1Llama 2-70B87.6%
- 2Claude 287.5%
- 3Claude 1.386.7%
- 4PaLM 2-L86.1%
- 5LLaMA-65B86.0%
- 6GPT 3.5 Turbo 110685.8%
- 7GPT 4 061384.8%
- 8Llama 2-34B84.6%
- 9LLaMA-33B83.8%
- 10DeepSeek V382.9%
- 11Llama 3.1 405B82.7%
- 12Mixtral 8x7B82.2%
- 13PaLM 2-M81.7%
- 14Palm 540B81.4%
- 15Deepseek v280.0%
36 of 36
| # | ||||
|---|---|---|---|---|
| 1 | 87.6% | 34.9 | 2023-07-18 | |
| 2 | 87.5% | – | 2023-07-11 | |
| 3 | 86.7% | – | 2023-04-18 | |
| 4 | 86.1% | – | 2023-05-17 | |
| 5 | 86.0% | – | 2023-02-24 | |
| 6 | 85.8% | 39.2 | 2023-11-06 | |
| 7 | 84.8% | 40.2 | 2023-06-13 | |
| 8 | 84.6% | – | 2023-07-18 | |
| 9 | 83.8% | – | 2023-02-24 | |
| 10 | 82.9% | 44.8 | 2024-12-26 | |
| 11 | 82.7% | 43.5 | 2024-07-23 | |
| 12 | 82.2% | 32.1 | 2023-12-11 | |
| 13 | 81.7% | – | 2023-05-17 | |
| 14 | 81.4% | – | 2022-04-04 | |
| 15 | 80.0% | – | 2024-05-07 | |
| 16 | 79.9% | – | 2023-03-15 | |
| 17 | 79.6% | 36.5 | 2023-07-18 | |
| 18 | 78.9% | – | – | |
| 19 | 78.7% | – | 2023-08-09 | |
| 20 | 77.9% | – | 2023-02-24 | |
| 21 | 75.2% | 32.1 | 2023-09-27 | |
| 22 | 75.2% | – | 2023-05-17 | |
| 23 | 73.9% | – | 2024-04-23 | |
| 24 | 73.7% | 34.2 | 2023-07-18 | |
| 25 | 73.6% | – | 2023-06-22 | |
| 26 | 72.3% | – | 2024-02-21 | |
| 27 | 71.9% | 42.7 | 2024-09-19 | |
| 28 | 71.2% | – | 2022-01-27 | |
| 29 | 71.0% | – | 2023-02-24 | |
| 30 | 67.7% | 35.6 | 2024-04-18 | |
| 31 | 64.6% | – | 2023-04-24 | |
| 32 | 64.0% | 36.2 | 2024-04-23 | |
| 33 | 61.6% | – | 2023-05-05 | |
| 34 | 58.1% | – | 2024-04-23 | |
| 35 | 53.2% | – | 2024-02-21 | |
| 36 | 45.2% | – | 2023-12-12 |
Cite as: BenchLeader, “TriviaQA leaderboard”, https://www.benchleader.com/benchmarks/triviaqa, data as of 19 Sept 2026.
TriviaQA: questions
- What does TriviaQA measure?
- Trivia questions with evidence documents. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads TriviaQA?
- Llama 2-70B leads TriviaQA with 87.6% as of 19 Sept 2026, ahead of Claude 2 at 87.5%.
- How many models have TriviaQA results?
- 36 model configurations have a TriviaQA result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
- Who runs TriviaQA and how often is it updated?
- TriviaQA is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does TriviaQA count toward the BenchLeader Index?
- No. TriviaQA is shown for reference but left out of the composite index.