BenchLeader

GPT-5.6 Luna vs Qwen3 8

Verdict
  • Qwen3 8 (max) leads on quality: 66.5 vs 61.1.
  • GPT-5.6 Luna (xhigh) is stronger in long context.
  • Qwen3 8 (max) is stronger in agents & tools, coding, composite, human preference, knowledge, multimodal, reasoning.
  • GPT-5.6 Luna (xhigh) is 5.9× cheaper ($0.450 vs $2.67 per 1M blended).
  • GPT-5.6 Luna (xhigh) streams 3.2× faster (124 vs 39 tokens per second).
MetricGPT-5.6 Luna (xhigh)Qwen3 8 (max)
BenchLeader Index61.166.5
Agents & tools score54.168.3
Coding score59.866.5
Composite score73.174.3
Human preference score64.868.2
Knowledge score59.163.5
Long context score67.366.6
Maths score67.067.0
Multimodal score61.667.4
Reasoning score60.176.3
Blended price $/M$0.450$2.67
Output speed124 tok/s39 tok/s
Time to first answer34.9 s55.1 s
Context window1.1M1M
SimpleBench46.8%
SciCode50.0%52.9%
ProofBench58.0%
Epoch Capabilities Index156.6
LMArena Text14531481
LMArena Hard Prompts14731503
LMArena Coding14981522
LMArena WebDev15191671
LMArena Vision12591315
LMArena Agent-0.43.3
LiveBench78.5%
LiveBench Reasoning88.2%
LiveBench Coding72.9%
LiveBench Agentic Coding64.7%
LiveBench Mathematics91.3%
LiveBench Data Analysis78.4%
LiveBench Language79.7%
LiveBench Instruction Following74.1%
AA Intelligence Index34.845.4
AA-LCR81.7%80.3%
MMMU-Pro78.5%82.8%
AA-Omniscience-10.812.0
GPQA Diamond (AA)89.5%92.8%
Humanity's Last Exam (AA)37.0%43.1%
SciCode (AA)50.5%53.2%
LiveCodeBench87.8%
MMLU-Pro88.6%
IOI68.9%
LegalBench83.6%
CorpFin65.8%
TaxEval75.5%
Terminal-Bench 2.1 (Vals)67.4%
SWE-bench (Vals)85.6%
GPQA Diamond (Vals)93.7%
Vals Index51.8
ARC-AGI-187.7%
ARC-AGI-247.6%
ARC-AGI-30.0%
CritPt20.6%20.0%
GDPval (AA)46.4%58.2%
τ²-Bench Banking (AA)28.7%51.3%
Code Migration24.0%
Excel Modeling Benchmark60.1%
Finance Agent v250.6%
Harvey's Legal Agent Benchmark10.4%
Legal Research Bench47.6%
MedCode40.7%
MedScribe85.0%
MMMU-Pro (Vals)88.0%
MortgageTax64.0%
MysteryMechanism23.9%
ProgramBench0.0%
Public Benefits Bench67.1%
SAGE51.3%
SkillsBench42.0%
Tax Agent Bench66.0%
Terminal-Bench 4.0 (Vals)24.8%
Terminal-Bench Science1.4%
Vals Multimodal Index65.4%
Vibe Code Bench 1-10012.8%
Vibe Code Bench v1.164.7%
LMArena Maths14771498
LMArena Creative Writing14111472
LMArena Instruction Following14421477
LMArena Multi-turn14581495
LMArena Longer Queries14501492
LMArena Document1462
DeepSWE56.9%
CursorBench57.7%

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

GPT-5.6 Luna vs Qwen3 8: questions

Is GPT-5.6 Luna better than Qwen3 8?
Qwen3 8 (max) leads on quality: 66.5 vs 61.1. The BenchLeader Index combines every independent quality benchmark; Qwen3 8 (max) is ahead overall as of 2026-09-19, but check the category scores for your use.
Is GPT-5.6 Luna better than Qwen3 8 for coding?
Qwen3 8 scores higher in coding (67 vs 60 on the category index, where 50 is average).
Is GPT-5.6 Luna better than Qwen3 8 for agentic tasks?
Qwen3 8 scores higher in agentic tasks (68 vs 54 on the category index, where 50 is average).
Which is cheaper, GPT-5.6 Luna or Qwen3 8?
GPT-5.6 Luna is cheaper: $0.450 against $2.67 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-5.6 Luna or Qwen3 8?
GPT-5.6 Luna streams faster: 124 against 39 output tokens per second.
Which has the larger context window?
GPT-5.6 Luna accepts more context: 1.1M against 1M tokens.