BenchLeader

GPT-5.6 Terra vs Kimi K2.6

Verdict
  • GPT-5.6 Terra (xhigh) leads on quality: 64.4 vs 60.7.
  • GPT-5.6 Terra (xhigh) is stronger in agents & tools, coding, composite, human preference, knowledge, maths.
  • Kimi K2.6 is stronger in instruction following, long context, multimodal, reasoning.
  • Kimi K2.6 is 2.6× cheaper ($1.71 vs $4.50 per 1M blended).
  • GPT-5.6 Terra (xhigh) streams 1.8× faster (79 vs 44 tokens per second).
MetricGPT-5.6 Terra (xhigh)Kimi K2.6
BenchLeader Index64.460.7
Agents & tools score69.845.1
Coding score63.760.4
Composite score77.368.6
Human preference score66.760.8
Instruction following score65.373.7
Knowledge score61.058.6
Long context score66.067.1
Maths score71.053.8
Multimodal score63.063.7
Reasoning score58.666.0
Blended price $/M$4.50$1.71
Output speed79 tok/s44 tok/s
Time to first answer33.6 s103.2 s
Context window1M262k
GPQA Diamond90.8%
FrontierMath Tiers 1–357.2%
FrontierMath Tier 425.6%
OTIS Mock AIME96.1%
SWE-bench Verified (Epoch)76.7%
SimpleQA Verified34.9%
SimpleBench48.9%
OSWorld-Verified 2.04.6%
SciCode51.6%53.5%
WeirdML55.9%
APEX-Agents18.9%
ProofBench74.0%16.0%
Epoch Capabilities Index151.0
LMArena Text14661461
LMArena Hard Prompts14911485
LMArena Coding15191514
LMArena WebDev15211509
LMArena Vision12691281
LMArena Agent1.5
AA Intelligence Index38.231.3
IFBench66.3%76.0%
AA-LCR79.0%81.0%
MMMU-Pro79.5%79.4%
AA-Omniscience-3.05.3
Terminal-Bench Hard62.9%43.9%
GPQA Diamond (AA)90.8%91.1%
Humanity's Last Exam (AA)41.9%37.5%
SciCode (AA)52.3%51.5%
τ²-Bench Telecom (AA)80.4%95.9%
LiveCodeBench85.9%86.8%
MMLU-Pro86.7%87.6%
IOI65.3%
LegalBench84.7%
CorpFin65.3%66.7%
TaxEval76.2%74.7%
Terminal-Bench 2.1 (Vals)53.6%
SWE-bench (Vals)76.2%
GPQA Diamond (Vals)90.9%89.1%
Vals Index43.5
HiL-Bench18.7%
EQ-Bench 41202
ARC-AGI-194.0%
ARC-AGI-274.2%
ARC-AGI-30.7%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

GPT-5.6 Terra vs Kimi K2.6: questions

Is GPT-5.6 Terra better than Kimi K2.6?
GPT-5.6 Terra (xhigh) leads on quality: 64.4 vs 60.7. The BenchLeader Index combines every independent quality benchmark; GPT-5.6 Terra (xhigh) is ahead overall as of 2026-09-10, but check the category scores for your use.
Is GPT-5.6 Terra better than Kimi K2.6 for coding?
GPT-5.6 Terra scores higher in coding (64 vs 60 on the category index, where 50 is average).
Is GPT-5.6 Terra better than Kimi K2.6 for agentic tasks?
GPT-5.6 Terra scores higher in agentic tasks (70 vs 45 on the category index, where 50 is average).
Which is cheaper, GPT-5.6 Terra or Kimi K2.6?
Kimi K2.6 is cheaper: $1.71 against $4.50 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-5.6 Terra or Kimi K2.6?
GPT-5.6 Terra streams faster: 79 against 44 output tokens per second.
Which has the larger context window?
GPT-5.6 Terra accepts more context: 1M against 262k tokens.