BenchLeader
OpenAIReasoning modelAuto-detected

GPT-5.6 Terra

Best configuration ranks #35 of 610 on the BenchLeader Index at 64.2 ±3.8 (xhigh reasoning effort). Last measured 8 Sept 2026. Released 9 Jul 2026.

Blended price
$4.50/M
$2.00 in · $12.00 out
Output speed
93 tok/s
First answer
33 s
first token 2.75 s
Context
1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index64
  2. Reasoning58
  3. Coding64
  4. Agents & tools70
  5. Maths71
  6. Knowledge61
  7. Instruction following65
  8. Human preference66
  9. Multimodal63
  10. Long context66
  11. Composite77

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning52.1#25783 tok/s0.95 s$0.0044Coding 50 · Composite 57 · Knowledge 53 · Long context 55 · Maths 47 · Multimodal 51 · Reasoning 57
low58.8#11683 tok/s1.95 s$0.0044Agents & tools 72 · Coding 57 · Composite 64 · Instruction following 59 · Knowledge 61 · Long context 62 · Maths 63 · Multimodal 61 · Reasoning 50
medium58.1#12980 tok/s2.07 s$0.0044Coding 57 · Composite 70 · Instruction following 62 · Knowledge 62 · Long context 63 · Multimodal 61 · Reasoning 50
high63.1#4983 tok/s3.08 s$0.0044Agents & tools 84 · Coding 64 · Composite 72 · Instruction following 63 · Knowledge 62 · Long context 65 · Multimodal 64 · Reasoning 59
xhighbest64.2#3593 tok/s33 s$0.0044Agents & tools 70 · Coding 64 · Composite 77 · Human preference 66 · Instruction following 65 · Knowledge 61 · Long context 66 · Maths 71 · Multimodal 63 · Reasoning 58
max58.0#13397 tok/s196 s$0.0048Agents & tools 45 · Coding 63 · Composite 60 · Knowledge 55 · Maths 72 · Reasoning 64
default62.3#5897 tok/s196 s$0.0044Agents & tools 84 · Coding 57 · Composite 82 · Human preference 55 · Instruction following 69 · Knowledge 64 · Long context 68 · Multimodal 65 · Reasoning 49

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
GPQA Diamond77.3%#12487.4%#6193.3%#14Epoch AI Benchmarking Hub
SimpleBench48.9%#4548.9%#45SimpleBench
LMArena Hard Prompts1491#40LMArena
LiveBench Reasoningnot in index90.6%#6LiveBench
GPQA Diamond (AA)not in index74.7%#24984.3%#14287.2%#10189.6%#7190.8%#5592.5%#35Artificial Analysis
Humanity's Last Exam (AA)not in index11.4%#23029.2%#11133.3%#9438.5%#6441.9%#4642.9%#34Artificial Analysis
GPQA Diamond (Vals)not in index90.9%#24Vals AI
Kagi LLM Benchmark51.3%#82Kagi LLM Benchmark
ARC-AGI-160.2%#10877.0%#8792.0%#4094.0%#2996.5%#13ARC Prize
ARC-AGI-218.8%#9537.5%#8267.1%#4374.2%#3183.9%#23ARC Prize
ARC-AGI-30.0%#380.1%#350.5%#220.7%#210.8%#20ARC Prize

Coding

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
SciCode44.6%#8949.2%#6549.6%#6450.1%#5951.6%#4953.9%#36SciCode
WeirdML78.3%#17WeirdML
FrontierCode41.3%#14Cognition
LMArena Coding1519#31LMArena
LMArena WebDev1521#33LMArena
LiveBench Codingnot in index78.3%#21LiveBench
SciCode (AA)not in index49.9%#7150.5%#6552.4%#5352.3%#5455.0%#34Artificial Analysis
LiveCodeBench85.9%#25Vals AI
IOI65.3%#7Vals AI
SWE-bench (Vals)not in index95.4%#5Vals AI

Agents & tools

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
Terminal-Bench21.5%#54Terminal-Bench
LMArena Agent1.5#22LMArena
LiveBench Agentic Codingnot in index55.0%#22LiveBench
Terminal-Bench Hard43.9%#3557.6%#1162.9%#257.6%#11Artificial Analysis
τ²-Bench Telecom (AA)not in index60.5%#17372.8%#14278.4%#12680.4%#11986.3%#87Artificial Analysis
Terminal-Bench 2.1 (Vals)77.5%#11Vals AI

Maths

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
FrontierMath Tiers 1–386.0%#6Epoch AI Benchmarking Hub
FrontierMath Tier 470.7%#15Epoch AI Benchmarking Hub
OTIS Mock AIME53.3%#16188.9%#6199.7%#8Epoch AI Benchmarking Hub
ProofBench74.0%#8Vals AI
LiveBench Mathematicsnot in index94.9%#9LiveBench

Knowledge

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
SimpleQA Verified43.2%#34Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index79.3%#9LiveBench
AA-Omniscience-23.2#194-6.8#122-5.1#117-3.5#109-3.0#1080.1#95Artificial Analysis
MMLU-Pro86.7%#39Vals AI
LegalBench85.1%#22Vals AI
CorpFin65.3%#29Vals AI
TaxEval76.2%#6Vals AI

Instruction following

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
LiveBench Languagenot in index82.9%#15LiveBench
IFBench59.7%#12562.2%#11764.4%#10666.3%#9671.2%#58Artificial Analysis

Human preference

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
LMArena Text1466#43LMArena
EQ-Bench 41234#11EQ-Bench

Multimodal

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
LMArena Vision1269#33LMArena
MMMU-Pro66.7%#13776.1%#6676.8%#6179.1%#4479.5%#4180.7%#30Artificial Analysis

Long context

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
AA-LCR58.7%#22471.3%#14974.0%#12077.7%#8779.0%#7183.0%#12Artificial Analysis

Composite

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
Epoch Capabilities Indexnot in index159.2#7Epoch AI Benchmarking Hub
LiveBench77.9%#14LiveBench
AA Intelligence Index22.3#14827.9#9632.8#6934.5#5638.2#4642.3#27Artificial Analysis
Vals Indexnot in index59.6#12Vals AI

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
OpenAI Flex70 tok/s2.25 s$1.00$6.001.1M
OpenAI Fast67 tok/s2.54 s$4.00$24.001.1M
Azure (EU)62 tok/s2.75 s$2.20$13.201.1M
OpenAI45 tok/s2.78 s$2.00$12.001.1M
Azure32 tok/s4.57 s$2.00$12.001.1M
Amazon Bedrock (US)31 tok/s1.34 s$2.20$13.201.1M
Azure (US)30 tok/s7.18 s$2.20$13.201.1M

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$4.13$8.25$12.38$16.50Jul 26Aug 26Sept 26
input outputnow $2.00 in · $12.00 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.220 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0044$0.003935.8 s
Summarise a 30-page report12,000 / 600$0.031$0.01539.0 s
Code edit6,000 / 1,500$0.030$0.02248.6 s
Agentic coding session60,000 / 4,000$0.168$0.0881.3 min
Structured extraction2,000 / 200$0.0064$0.003734.7 s

See also

Data as of 9 Sept 2026. Compare these configurations.