BenchLeader

Terminal-Bench Science

Expert-authored scientific research workflows in a terminal. Run by Vals AI.

As of 19 Sept 2026, GPT-6 Astra leads Terminal-Bench Science on BenchLeader with 65.7%, ahead of Claude Fable 5.1 at 34.3%, across 25 model configurations with a published result.

Published by
Vals AI
Category
Agents & tools
Index weight
Reference only
Models
25
Data as of
19 Sept 2026

Vals AI (vals.ai).

What the test looks like

Scientists wrote multi-step research workflows, from data processing to analysis, for an agent to complete in a shell.

How it is scored

Percent of workflows completed, run by Vals AI.

What to keep in mind

Domain knowledge and tooling both matter; few models measured.

25 of 25
#
1GPT-6 AstramaxOpenAI65.7%71.82026-09-16
2Claude Fable 5.1Anthropic34.3%64.52026-09-16
3Claude Opus 5Anthropic22.9%66.72026-09-16
4Muse Spark 1.3maxMeta14.3%69.32026-09-16
5GPT-5.6 SolmaxOpenAI12.9%68.82026-09-16
6Claude Fable 5Anthropic12.9%68.32026-09-16
7GPT-5.6 TerramaxOpenAI10.0%65.02026-09-16
8Claude Opus 4.8Anthropic10.0%61.92026-09-16
9Muse Spark 1.3xhighMeta8.6%68.02026-09-16
10Gemini 3.8 FlashhighGoogle8.6%64.52026-09-16
11Gemini 3.7 FlashhighGoogle5.7%64.62026-09-16
12GPT-5.6 LunamaxOpenAI5.7%60.22026-09-16
13GLM-5.3maxZhipu AIopen ↗4.3%65.82026-09-16
14Grok 4.6highxAI4.3%64.52026-09-16
15Gemini 3.5 FlashhighGoogle4.3%63.62026-09-16
16DeepSeek V4.1 FlashhighDeepSeekopen ↗4.3%2026-09-16
17Kimi K3maxMoonshot AIopen ↗2.9%67.22026-09-16
18Claude Sonnet 5Anthropic2.9%57.22026-09-16
19Qwen3 8maxAlibaba1.4%66.52026-09-16
20Gemini 3.6 FlashhighGoogle1.4%61.72026-09-16
21Qwen3.8 27BxhighAlibabaopen ↗1.4%59.22026-09-16
22DeepSeek V4 Pro 0813maxDeepSeekopen ↗0.0%56.02026-09-16
23GLM-5.3-FlashmaxZhipu AIopen ↗0.0%55.32026-09-16
24DeepSeek V4 Flash 0731highDeepSeekopen ↗0.0%55.02026-09-16
25Mercury 2.5highInception0.0%2026-09-16

Cite as: BenchLeader, “Terminal-Bench Science leaderboard”, https://www.benchleader.com/benchmarks/vals_terminal_bench_science, data as of 19 Sept 2026.

Terminal-Bench Science: questions

What does Terminal-Bench Science measure?
Scientists wrote multi-step research workflows, from data processing to analysis, for an agent to complete in a shell. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads Terminal-Bench Science?
GPT-6 Astra leads Terminal-Bench Science with 65.7% as of 19 Sept 2026, ahead of Claude Fable 5.1 at 34.3%.
How many models have Terminal-Bench Science results?
25 model configurations have a Terminal-Bench Science result on BenchLeader, all taken from Vals AI.
Who runs Terminal-Bench Science and how often is it updated?
Terminal-Bench Science is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
Does Terminal-Bench Science count toward the BenchLeader Index?
No. Terminal-Bench Science is shown for reference but left out of the composite index.