BenchLeader
OpenAIReasoning modelAuto-detected

GPT-6 Astra

Best configuration ranks #1 of 610 on the BenchLeader Index at 70.6 ±3.6 (max reasoning effort). Last measured 8 Sept 2026. Released 3 Sept 2026.

Blended price
$22.00/M
$11.00 in · $55.00 out
Output speed
54 tok/s
First answer
329 s
first token 7.00 s
Context
1.1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index71
  2. Reasoning71
  3. Coding79
  4. Agents & tools68
  5. Maths77
  6. Knowledge80
  7. Composite73

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning64.4#3354 tok/s329 s$0.021Coding 62 · Composite 86 · Knowledge 77 · Long context 62 · Maths 72 · Multimodal 67 · Reasoning 60
low65.8#2447 tok/s2.69 s$0.019Agents & tools 56 · Coding 63 · Composite 87 · Knowledge 83 · Long context 67 · Maths 74 · Multimodal 69 · Reasoning 65
medium67.4#1544 tok/s5.42 s$0.019Agents & tools 58 · Coding 63 · Composite 92 · Knowledge 84 · Long context 66 · Maths 79 · Multimodal 70 · Reasoning 69
high69.5#949 tok/s93 s$0.019Agents & tools 61 · Coding 72 · Composite 93 · Knowledge 85 · Long context 67 · Maths 79 · Multimodal 71 · Reasoning 71
xhigh68.6#1151 tok/s209 s$0.019Agents & tools 61 · Coding 65 · Composite 95 · Knowledge 85 · Long context 67 · Maths 79 · Multimodal 71 · Reasoning 71
maxbest70.6#154 tok/s329 s$0.021Agents & tools 68 · Coding 79 · Composite 73 · Knowledge 80 · Maths 77 · Reasoning 71
default69.9#754 tok/s329 s$0.019Agents & tools 75 · Composite 95 · Knowledge 72 · Long context 67 · Maths 85 · Multimodal 72

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
GPQA Diamond95.8%#1Epoch AI Benchmarking Hub
LiveBench Reasoningnot in index92.7%#1LiveBench
GPQA Diamond (AA)not in index89.5%#7493.1%#2493.9%#1095.0%#496.3%#196.1%#2Artificial Analysis
Humanity's Last Exam (AA)not in index37.1%#7149.2%#1552.7%#1253.1%#1054.6%#754.7%#6Artificial Analysis
ARC-AGI-186.0%#7096.5%#1397.5%#698.5%#198.5%#197.5%#6ARC Prize
ARC-AGI-259.6%#6085.4%#1692.1%#492.1%#493.3%#295.0%#1ARC Prize
ARC-AGI-335.2%#1117.5%#1338.6%#1054.8%#959.3%#862.7%#7ARC Prize

Coding

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
SciCode53.5%#4154.0%#3454.2%#3355.4%#2455.7%#2156.5%#13SciCode
WeirdML92.9%#2WeirdML
FrontierCode53.3%#3Cognition
LMArena WebDev1796#1LMArena
LiveBench Codingnot in index80.4%#13LiveBench
SciCode (AA)not in index53.5%#4754.0%#4254.2%#4155.4%#3155.7%#2856.5%#20Artificial Analysis

Agents & tools

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
Terminal-Bench50.6%#2554.2%#2157.9%#1657.9%#1658.2%#15Terminal-Bench
APEX-Agents46.7%#2Mercor
LMArena Agent12.5#2LMArena
LiveBench Agentic Codingnot in index57.3%#17LiveBench
Terminal-Bench 2.1 (Vals)87.3%#1Vals AI

Maths

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
FrontierMath Tiers 1–393.7%#1Epoch AI Benchmarking Hub
FrontierMath Tier 482.9%#887.8%#697.6%#497.6%#197.6%#197.6%#1Epoch AI Benchmarking Hub
OTIS Mock AIME100.0%#1Epoch AI Benchmarking Hub
ProofBench99.0%#2Vals AI
LiveBench Mathematicsnot in index96.8%#2LiveBench

Knowledge

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
SimpleQA Verified75.6%#1Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index83.0%#1LiveBench
AA-Omniscience26.6#2740.5#942.2#743.7#143.4#343.4#4Artificial Analysis
PRBench Finance47.5%#13Scale AI SEAL
PRBench Legal48.4%#13Scale AI SEAL

Instruction following

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
LiveBench Languagenot in index89.4%#3LiveBench

Multimodal

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
MMMU-Pro82.0%#2284.6%#1085.1%#686.4%#286.2%#386.9%#1Artificial Analysis

Long context

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
AA-LCR70.7%#15480.0%#5179.7%#5980.0%#5180.0%#5180.7%#41Artificial Analysis

Composite

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
Epoch Capabilities Indexnot in index166.6#1Epoch AI Benchmarking Hub
LiveBench82.2%#3LiveBench
AA Intelligence Index45.2#1746.0#1649.7#951.0#652.5#452.8#3Artificial Analysis
Vals Indexnot in index66.6#3Vals AI

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
OpenAI Flex55 tok/s3.69 s$5.00$25.001.1M
OpenAI31 tok/s5.03 s$10.00$50.001.1M
OpenAI Fast27 tok/s7.00 s$20.00$100.001.1M
Azure9 tok/s15 s$10.00$50.001.1M
Azure (US)3 tok/s66 s$11.00$55.001.1M

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $1.10 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.021$0.0185.6 min
Summarise a 30-page report12,000 / 600$0.165$0.0765.7 min
Code edit6,000 / 1,500$0.148$0.1045.9 min
Agentic coding session60,000 / 4,000$0.880$0.4346.7 min
Structured extraction2,000 / 200$0.033$0.0185.5 min

See also

Data as of 9 Sept 2026. Compare these configurations.