OpenBookQA
Elementary science questions with an open book of facts. A 2018 benchmark, saturated for current models; kept for history.
As of 19 Sept 2026, Phi 3 Mini leads OpenBookQA on BenchLeader with 88.0%, ahead of Phi 3 Small at 88.0%, across 41 model configurations with a published result.
- Published by
- Epoch AI Benchmarking Hub
- Category
- Knowledge
- Index weight
- Reference only
- Models
- 41
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
Elementary science questions with an open book of facts.
How it is scored
Accuracy, as compiled by Epoch AI from published results.
What to keep in mind
Saturated: frontier models score near the ceiling, so it separates only older or smaller models, and results are largely self-reported by developers.
- 1Phi 3 Mini88.0%
- 2Phi 3 Small88.0%
- 3phi-3-medium 14B (medium)87.4%
- 4GPT 3.5 Turbo 110686.0%
- 5Mixtral 8x7B85.8%
- 6Llama 3-8B82.6%
- 7Mistral 7B v0.379.8%
- 8Gemma 7B78.6%
- 9Phi-273.6%
- 10Palm 540B68.0%
- 11InstructGPT 175B65.4%
- 12Falcon-180B64.2%
- 13Llama 2-70B60.2%
- 14LLaMA-65B60.2%
- 15Llama 2-7B58.6%
| # | ||||
|---|---|---|---|---|
| 1 | 88.0% | 36.2 | 2024-04-23 | |
| 2 | 88.0% | – | 2024-04-23 | |
| 3 | 87.4% | – | 2024-04-23 | |
| 4 | 86.0% | 39.2 | 2023-11-06 | |
| 5 | 85.8% | 32.1 | 2023-12-11 | |
| 6 | 82.6% | 35.6 | 2024-04-18 | |
| 7 | 79.8% | 32.1 | 2023-09-27 | |
| 8 | 78.6% | – | 2024-02-21 | |
| 9 | 73.6% | – | 2023-12-12 | |
| 10 | 68.0% | – | 2022-04-04 | |
| 11 | 65.4% | – | 2022-01-27 | |
| 12 | 64.2% | – | 2023-09-06 | |
| 13 | 60.2% | 34.9 | 2023-07-18 | |
| 14 | 60.2% | – | 2023-02-24 | |
| 15 | 58.6% | 34.2 | 2023-07-18 | |
| 16 | 58.6% | – | 2023-02-24 | |
| 17 | 58.2% | – | 2023-07-18 | |
| 18 | 57.4% | – | 2023-05-17 | |
| 19 | 57.2% | – | 2023-02-24 | |
| 20 | 57.0% | 36.5 | 2023-07-18 | |
| 21 | 56.6% | – | 2023-03-15 | |
| 22 | 56.4% | – | 2023-02-24 | |
| 23 | 56.2% | – | 2023-05-17 | |
| 24 | 52.0% | – | 2023-06-22 | |
| 25 | 51.6% | – | 2023-04-24 | |
| 26 | 51.4% | – | 2023-05-05 | |
| 27 | 50.4% | – | 2022-04-04 | |
| 28 | 40.2% | – | 2023-06-27 | |
| 29 | RedPajama-INCITE-7B-BaseUnknown | 40.0% | – | 2023-05-04 |
| 30 | opt-13bUnknown | 39.8% | – | 2022-05-11 |
| 31 | 39.2% | – | 2023-04-11 | |
| 32 | open_llama_7bUnknown | 39.0% | – | 2023-06-07 |
| 33 | GPT-NeoX-20BEleutherAI | 38.8% | – | 2022-04-07 |
| 34 | GPT-J-6BEleutherAI | 38.2% | – | 2021-08-05 |
| 35 | 37.2% | – | 2023-09-11 | |
| 36 | 35.8% | – | 2023-03-20 | |
| 37 | 33.0% | – | 2023-04-12 | |
| 38 | 32.4% | – | 2023-04-19 | |
| 39 | 24.0% | – | 2022-05-11 | |
| 40 | GPT-Neo-2.7BEleutherAI | 23.2% | – | 2021-03-30 |
| 41 | 22.4% | – | 2019-11-05 |
Cite as: BenchLeader, “OpenBookQA leaderboard”, https://www.benchleader.com/benchmarks/openbookqa, data as of 19 Sept 2026.
OpenBookQA: questions
- What does OpenBookQA measure?
- Elementary science questions with an open book of facts. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads OpenBookQA?
- Phi 3 Mini leads OpenBookQA with 88.0% as of 19 Sept 2026, ahead of Phi 3 Small at 88.0%.
- How many models have OpenBookQA results?
- 41 model configurations have a OpenBookQA result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
- Who runs OpenBookQA and how often is it updated?
- OpenBookQA is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does OpenBookQA count toward the BenchLeader Index?
- No. OpenBookQA is shown for reference but left out of the composite index.