BenchLeader

GPT-6.1 Sol vs GPT-6 Astra

Verdict
  • GPT-6.1 Sol (max) and GPT-6 Astra (max) are level on quality (69.3 vs 70.2).
  • GPT-6.1 Sol (max) is stronger in human preference, long context, maths.
  • GPT-6 Astra (max) is stronger in agents & tools, coding, composite, knowledge, multimodal, reasoning.
  • GPT-6.1 Sol (max) is 5.0× cheaper ($4.00 vs $20.00 per 1M blended).
MetricGPT-6.1 Sol (max)GPT-6 Astra (max)
BenchLeader Index69.370.2
Agents & tools score63.968.8
Coding score72.474.4
Composite score79.080.3
Human preference score68.167.0
Knowledge score78.879.9
Long context score66.865.6
Maths score72.572.2
Multimodal score66.266.3
Reasoning score76.476.7
Blended price $/M$4.00$20.00
Output speed56 tok/s47 tok/s
Time to first answer326.9 s383.6 s
Context window1.1M1.1M
GPQA Diamond95.4%95.8%
FrontierMath Tiers 1–393.7%93.7%
FrontierMath Tier 4100.0%97.6%
OTIS Mock AIME100.0%100.0%
SimpleQA Verified73.9%75.6%
Terminal-Bench58.2%58.2%
SciCode54.2%56.5%
WeirdML–93.3%
APEX-Agents60.0%–
FrontierCode–53.3%
LMArena Text14841475
LMArena Hard Prompts15071500
LMArena Coding15451542
LMArena WebDev17551786
LMArena Vision12881281
LMArena Agent11.713.1
LiveBench81.6%82.2%
LiveBench Reasoning92.6%92.7%
LiveBench Coding80.4%80.4%
LiveBench Agentic Coding54.5%57.3%
LiveBench Mathematics96.8%96.8%
LiveBench Data Analysis82.7%83.0%
LiveBench Language90.1%89.4%
LiveBench Instruction Following74.2%75.6%
AA Intelligence Index v4.3.251.852.7
AA-LCR83.0%80.7%
MMMU-Pro86.0%86.9%
AA-Omniscience41.543.4
GPQA Diamond (AA)–96.1%
Humanity's Last Exam (AA)52.9%54.7%
SciCode (AA)54.2%56.5%
IOI96.9%100.0%
Terminal-Bench 2.1 (Vals)–87.3%
Vals Index61.163.1
ARC-AGI-196.5%97.5%
ARC-AGI-294.2%95.0%
ARC-AGI-396.2%98.6%
CritPt31.7%31.7%
GDPval-AA v2.153.8%52.1%
τ³-Banking (AA)–41.4%
ITBench SRE (AA)–48.6%
Analyst Agent (AA)50.0%51.3%
BioMysteryBench79.6%79.3%
Code Migration65.1%67.7%
CUA-bench–19.2%
CyberBench39.3%41.1%
Excel Modeling Benchmark70.8%71.7%
Finance Agent v252.0%53.5%
Harvey's Legal Agent Benchmark5.4%5.4%
Legal Research Bench38.5%39.4%
MedCode48.8%48.5%
MedScribe86.5%87.9%
MysteryMechanism46.4%53.1%
ProgramBench–5.5%
Public Benefits Bench59.3%–
SAGE46.5%46.4%
SREBench50.8%56.9%
Tax Agent Bench62.3%63.3%
Terminal-Bench 4.0 (Vals)55.0%59.6%
Terminal-Bench Science–62.9%
Time Horizon Index: KSP–90.5%
Vibe Code Bench 1-100–27.6%
Vibe Code Bench v1.188.9%89.6%
LMArena Maths14881486
LMArena Creative Writing14621448
LMArena Instruction Following14881469
LMArena Multi-turn14871484
LMArena Longer Queries14961484
LMArena Document–1473
Chess Puzzles61.0%72.0%
EBR-bench54.3%76.2%
Mystery Game Puzzles80.0%84.0%
BALROG–68.3%
DeepSWE v1.1–73.2%
ALE-Bench–2951.3
GDP.pdf–34.2%
FrontierSWE–65.5%
Terminal-Bench 4.0 (AA)56.1%59.1%
Terminal-Bench 2.1 (AA)–88.4%
AutomationBench64.9%68.5%
GDP.pdf31.0%31.0%
MLCR33.9%35.0%
Harvey LAB6.9%8.6%
AA-Omniscience: accuracy62.1%62.6%
AA-Omniscience: non-hallucination45.7%48.7%
AA-Briefcase v1.115571570

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

GPT-6.1 Sol vs GPT-6 Astra: questions

Is GPT-6.1 Sol better than GPT-6 Astra?
GPT-6.1 Sol (max) and GPT-6 Astra (max) are level on quality (69.3 vs 70.2). The BenchLeader Index combines every independent quality benchmark; GPT-6 Astra (max) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is GPT-6.1 Sol better than GPT-6 Astra for coding?
GPT-6 Astra scores higher in coding (74 vs 72 on the category index, where 50 is average).
Is GPT-6.1 Sol better than GPT-6 Astra for agentic tasks?
GPT-6 Astra scores higher in agentic tasks (69 vs 64 on the category index, where 50 is average).
Which is cheaper, GPT-6.1 Sol or GPT-6 Astra?
GPT-6.1 Sol is cheaper: $4.00 against $20.00 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-6.1 Sol or GPT-6 Astra?
GPT-6.1 Sol streams faster: 56 against 47 output tokens per second.
Which has the larger context window?
Both accept 1.1M tokens of context.