BenchLeader

GPT-5.6 Terra vs o3

Verdict
  • GPT-5.6 Terra (xhigh) leads on quality: 64.1 vs 61.6.
  • GPT-5.6 Terra (xhigh) is stronger in composite, human preference, instruction following, knowledge, long context, multimodal.
  • o3 is stronger in agents & tools, coding, maths, reasoning.
  • o3 is 1.3× cheaper ($3.50 vs $4.50 per 1M blended).
  • o3 streams 1.4× faster (107 vs 76 tokens per second).
MetricGPT-5.6 Terra (xhigh)o3
BenchLeader Index64.161.6
Agents & tools score69.570.3
Coding score61.562.0
Composite score77.454.5
Human preference score66.562.3
Instruction following score65.365.0
Knowledge score61.058.0
Long context score66.063.8
Maths score71.078.6
Multimodal score63.057.9
Reasoning score58.661.5
Blended price $/M$4.50$3.50
Output speed76 tok/s107 tok/s
Time to first answer30.9 s6.3 s
Context window1.1M200k
SimpleBench48.9%
SciCode51.6%
ProofBench74.0%
Epoch Capabilities Index146.9
LMArena Text14661432
LMArena Hard Prompts14891441
LMArena Coding15181460
LMArena WebDev1521
LMArena Vision12691214
LMArena Agent1.5
AA Intelligence Index38.220.2
IFBench66.3%71.4%
AA-LCR79.0%74.7%
MMMU-Pro79.5%70.1%
AA-Omniscience-3.0-15.6
Terminal-Bench Hard62.9%37.1%
GPQA Diamond (AA)90.8%82.7%
Humanity's Last Exam (AA)41.9%20.1%
SciCode (AA)52.3%
τ²-Bench Telecom (AA)80.4%80.7%
LiveCodeBench85.9%
MMLU-Pro86.7%
CorpFin65.3%
TaxEval76.2%
GPQA Diamond (Vals)90.9%
PRBench Finance47.7%
PRBench Legal48.6%
MMMU (validation)82.9%
MMMU-Pro (official)76.4%
Kagi LLM Benchmark67.6%
IFEval (HELM)86.9%
Omni-MATH (HELM)71.4%
WildBench (HELM)86.1%
MMLU-Pro (HELM)85.9%
GPQA Diamond (HELM)75.3%
HELM Capabilities mean81.1%
Aider Polyglot81.3%
SWE-bench Verified (bash only)58.4%
SWE-bench Verified (any scaffold)58.4%
ARC-AGI-194.0%
ARC-AGI-274.2%
ARC-AGI-30.7%
BFCL Overall63.0%

Data as of 2026-09-12. Best configuration of each model; every score links to its source on the model pages.

GPT-5.6 Terra vs o3: questions

Is GPT-5.6 Terra better than o3?
GPT-5.6 Terra (xhigh) leads on quality: 64.1 vs 61.6. The BenchLeader Index combines every independent quality benchmark; GPT-5.6 Terra (xhigh) is ahead overall as of 2026-09-12, but check the category scores for your use.
Is GPT-5.6 Terra better than o3 for coding?
o3 scores higher in coding (62 vs 62 on the category index, where 50 is average).
Is GPT-5.6 Terra better than o3 for agentic tasks?
o3 scores higher in agentic tasks (70 vs 70 on the category index, where 50 is average).
Which is cheaper, GPT-5.6 Terra or o3?
o3 is cheaper: $3.50 against $4.50 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-5.6 Terra or o3?
o3 streams faster: 107 against 76 output tokens per second.
Which has the larger context window?
GPT-5.6 Terra accepts more context: 1.1M against 200k tokens.