ScienceQA
Science questions with diagrams. A 2022 benchmark, saturated for current models; kept for history.
As of 19 Sept 2026, Phi 3.5 Vision leads ScienceQA on BenchLeader with 91.3%, ahead of GPT-4o at 88.5%, across 23 model configurations with a published result.
- Published by
- Epoch AI Benchmarking Hub
- Category
- Multimodal
- Index weight
- Reference only
- Models
- 23
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
Science questions with diagrams.
How it is scored
Accuracy, as compiled by Epoch AI from published results.
What to keep in mind
Saturated: frontier models score near the ceiling, so it separates only older or smaller models, and results are largely self-reported by developers.
- 1Phi 3.5 Vision91.3%
- 2GPT-4o88.5%
- 3Gemini 1.0 Pro Vision79.7%
- 4falcon-11B-vlm74.9%
- 5Blip2 Opt 2 7B74.2%
- 6InstructGPT 175B74.0%
- 7llama3-llava-next-8b73.7%
- 8llava-v1.6-vicuna-13b73.6%
- 9llava-v1.6-mistral-7b72.8%
- 10MM1-7B-Chat72.6%
- 11Claude 3 Haiku72.0%
- 12llava-v1.6-vicuna-7b70.6%
- 13InternVL-Chat-ViT-6B-Vicuna-13B70.1%
- 14MM1-3B-Chat69.4%
- 15Qwen Vl68.2%
23 of 23
| # | ||||
|---|---|---|---|---|
| 1 | 91.3% | – | 2024-08-16 | |
| 2 | 88.5% | 44.3 | 2024-05-13 | |
| 3 | 79.7% | – | 2024-01-04 | |
| 4 | 74.9% | – | 2024-05-21 | |
| 5 | Blip2 Opt 2 7BSalesforce Research | 74.2% | – | 2023-02-06 |
| 6 | 74.0% | – | 2022-01-27 | |
| 7 | 73.7% | – | 2024-04-20 | |
| 8 | 73.6% | – | 2024-01-31 | |
| 9 | 72.8% | – | 2024-01-31 | |
| 10 | MM1-7B-ChatUnknown | 72.6% | – | 2024-03-14 |
| 11 | 72.0% | 37.6 | 2024-03-07 | |
| 12 | 70.6% | – | 2024-01-31 | |
| 13 | 70.1% | – | 2024-08-16 | |
| 14 | MM1-3B-ChatUnknown | 69.4% | – | 2024-03-14 |
| 15 | 68.2% | – | 2023-08-20 | |
| 16 | 66.8% | – | 2023-10-05 | |
| 17 | 66.2% | – | 2023-12-25 | |
| 18 | instructblip-vicuna-13bUnknown | 63.1% | – | 2023-12-25 |
| 19 | instructblip-vicuna-7bUnknown | 60.5% | – | 2023-05-22 |
| 20 | 55.8% | 36.5 | 2023-07-18 | |
| 21 | 43.3% | – | 2023-02-24 | |
| 22 | 43.1% | 34.2 | 2023-07-18 | |
| 23 | 36.2% | – | 2023-02-24 |
Cite as: BenchLeader, “ScienceQA leaderboard”, https://www.benchleader.com/benchmarks/scienceqa, data as of 19 Sept 2026.
ScienceQA: questions
- What does ScienceQA measure?
- Science questions with diagrams. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads ScienceQA?
- Phi 3.5 Vision leads ScienceQA with 91.3% as of 19 Sept 2026, ahead of GPT-4o at 88.5%.
- How many models have ScienceQA results?
- 23 model configurations have a ScienceQA result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
- Who runs ScienceQA and how often is it updated?
- ScienceQA is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does ScienceQA count toward the BenchLeader Index?
- No. ScienceQA is shown for reference but left out of the composite index.