Terminal-Bench Science
Expert-authored scientific research workflows in a terminal. Run by Vals AI.
As of 19 Sept 2026, GPT-6 Astra leads Terminal-Bench Science on BenchLeader with 65.7%, ahead of Claude Fable 5.1 at 34.3%, across 25 model configurations with a published result.
- Published by
- Vals AI
- Category
- Agents & tools
- Index weight
- Reference only
- Models
- 25
- Data as of
- 19 Sept 2026
Vals AI (vals.ai).
What the test looks like
Scientists wrote multi-step research workflows, from data processing to analysis, for an agent to complete in a shell.
How it is scored
Percent of workflows completed, run by Vals AI.
What to keep in mind
Domain knowledge and tooling both matter; few models measured.
- 1GPT-6 Astra (max)65.7%
- 2Claude Fable 5.134.3%
- 3Claude Opus 522.9%
- 4Muse Spark 1.3 (max)14.3%
- 5GPT-5.6 Sol (max)12.9%
- 6Claude Fable 512.9%
- 7GPT-5.6 Terra (max)10.0%
- 8Claude Opus 4.810.0%
- 9Muse Spark 1.3 (xhigh)8.6%
- 10Gemini 3.8 Flash (high)8.6%
- 11Gemini 3.7 Flash (high)5.7%
- 12GPT-5.6 Luna (max)5.7%
- 13GLM-5.3 (max)4.3%
- 14Grok 4.6 (high)4.3%
- 15Gemini 3.5 Flash (high)4.3%
25 of 25
| # | ||||
|---|---|---|---|---|
| 1 | 65.7% | 71.8 | 2026-09-16 | |
| 2 | 34.3% | 64.5 | 2026-09-16 | |
| 3 | 22.9% | 66.7 | 2026-09-16 | |
| 4 | 14.3% | 69.3 | 2026-09-16 | |
| 5 | 12.9% | 68.8 | 2026-09-16 | |
| 6 | 12.9% | 68.3 | 2026-09-16 | |
| 7 | 10.0% | 65.0 | 2026-09-16 | |
| 8 | 10.0% | 61.9 | 2026-09-16 | |
| 9 | 8.6% | 68.0 | 2026-09-16 | |
| 10 | 8.6% | 64.5 | 2026-09-16 | |
| 11 | 5.7% | 64.6 | 2026-09-16 | |
| 12 | 5.7% | 60.2 | 2026-09-16 | |
| 13 | 4.3% | 65.8 | 2026-09-16 | |
| 14 | 4.3% | 64.5 | 2026-09-16 | |
| 15 | 4.3% | 63.6 | 2026-09-16 | |
| 16 | 4.3% | – | 2026-09-16 | |
| 17 | 2.9% | 67.2 | 2026-09-16 | |
| 18 | 2.9% | 57.2 | 2026-09-16 | |
| 19 | 1.4% | 66.5 | 2026-09-16 | |
| 20 | 1.4% | 61.7 | 2026-09-16 | |
| 21 | 1.4% | 59.2 | 2026-09-16 | |
| 22 | 0.0% | 56.0 | 2026-09-16 | |
| 23 | 0.0% | 55.3 | 2026-09-16 | |
| 24 | 0.0% | 55.0 | 2026-09-16 | |
| 25 | 0.0% | – | 2026-09-16 |
Cite as: BenchLeader, “Terminal-Bench Science leaderboard”, https://www.benchleader.com/benchmarks/vals_terminal_bench_science, data as of 19 Sept 2026.
Terminal-Bench Science: questions
- What does Terminal-Bench Science measure?
- Scientists wrote multi-step research workflows, from data processing to analysis, for an agent to complete in a shell. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads Terminal-Bench Science?
- GPT-6 Astra leads Terminal-Bench Science with 65.7% as of 19 Sept 2026, ahead of Claude Fable 5.1 at 34.3%.
- How many models have Terminal-Bench Science results?
- 25 model configurations have a Terminal-Bench Science result on BenchLeader, all taken from Vals AI.
- Who runs Terminal-Bench Science and how often is it updated?
- Terminal-Bench Science is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does Terminal-Bench Science count toward the BenchLeader Index?
- No. Terminal-Bench Science is shown for reference but left out of the composite index.