BenchLeader

Qwen3.8 Max vs Step 5 Preview

Verdict
  • Qwen3.8 Max (max) leads on quality: 64.2 vs 62.4.
  • Qwen3.8 Max (max) is stronger in agents & tools, human preference, maths, multimodal.
  • Step 5 Preview is stronger in coding, composite, knowledge, long context.
  • Step 5 Preview is 2.1× cheaper ($1.43 vs $3.00 per 1M blended).
  • Step 5 Preview streams 2.4× faster (87 vs 37 tokens per second).
MetricQwen3.8 Max (max)Step 5 Preview
BenchLeader Index64.262.4
Agents & tools score58.247.2
Coding score64.665.6
Composite score70.779.1
Human preference score68.063.6
Knowledge score62.769.1
Long context score65.469.5
Maths score65.161.1
Multimodal score66.260.0
Reasoning score72.272.2
Blended price $/M$3.00$1.43
Output speed37 tok/s87 tok/s
Time to first answer58.9 s25.9 s
Context window1M1M
Terminal-Bench27.0%–
SciCode53.2%58.9%
APEX-Agents63.3%–
ProofBench58.0%42.0%
Epoch Capabilities Index156.4–
LMArena Text14831447
LMArena Hard Prompts15041479
LMArena Coding15241503
LMArena WebDev1672–
LMArena Vision13141267
LMArena Agent2.30
LiveBench78.5%–
LiveBench Reasoning88.2%–
LiveBench Coding72.9%–
LiveBench Agentic Coding64.7%–
LiveBench Mathematics91.3%–
LiveBench Data Analysis78.4%–
LiveBench Language79.7%–
LiveBench Instruction Following74.1%–
AA Intelligence Index v4.3.245.443.7
AA-LCR80.3%88.3%
MMMU-Pro82.8%76.4%
AA-Omniscience12.016.4
GPQA Diamond (AA)92.8%–
Humanity's Last Exam (AA)43.1%46.5%
SciCode (AA)53.2%58.9%
LiveCodeBench87.8%–
MMLU-Pro88.6%–
IOI68.9%–
LegalBench83.6%–
CorpFin65.8%–
TaxEval75.5%–
Terminal-Bench 2.1 (Vals)67.4%–
SWE-bench (Vals)85.6%–
GPQA Diamond (Vals)93.7%–
Vals Index48.3–
CritPt20.0%20.9%
GDPval-AA v2.158.6%54.2%
τ³-Banking (AA)51.3%–
ITBench SRE (AA)40.3%55.6%
Analyst Agent (AA)45.0%35.0%
APEX-Agents (AA)42.4%38.0%
Code Migration24.0%–
CyberBench28.6%–
Excel Modeling Benchmark60.1%–
Finance Agent v250.6%–
Harvey's Legal Agent Benchmark10.4%–
Legal Research Bench47.6%–
MedCode40.7%–
MedScribe85.0%–
MMMU-Pro (Vals)88.0%–
MortgageTax64.0%–
MysteryMechanism23.9%–
ProgramBench0.0%–
Public Benefits Bench67.1%–
SAGE51.3%–
SkillsBench42.0%–
Tax Agent Bench66.0%–
Terminal-Bench 4.0 (Vals)34.3%–
Terminal-Bench Science1.4%–
Vals Multimodal Index65.4%–
Vibe Code Bench 1-10012.8%–
Vibe Code Bench v1.164.7%–
LMArena Maths14971478
LMArena Creative Writing14701406
LMArena Instruction Following14741449
LMArena Multi-turn14921453
LMArena Longer Queries14921462
Terminal-Bench 4.0 (AA)38.9%33.3%
Terminal-Bench 2.1 (AA)88.8%–
AutomationBench56.2%51.0%
GDP.pdf22.8%14.8%
MLCR20.0%16.7%
EnterpriseOps-Gym47.6%47.2%
AA-Omniscience: accuracy31.9%41.5%
AA-Omniscience: non-hallucination71.2%57.0%
AA-Briefcase v1.116171425

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

Qwen3.8 Max vs Step 5 Preview: questions

Is Qwen3.8 Max better than Step 5 Preview?
Qwen3.8 Max (max) leads on quality: 64.2 vs 62.4. The BenchLeader Index combines every independent quality benchmark; Qwen3.8 Max (max) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is Qwen3.8 Max better than Step 5 Preview for coding?
Step 5 Preview scores higher in coding (66 vs 65 on the category index, where 50 is average).
Is Qwen3.8 Max better than Step 5 Preview for agentic tasks?
Qwen3.8 Max scores higher in agentic tasks (58 vs 47 on the category index, where 50 is average).
Which is cheaper, Qwen3.8 Max or Step 5 Preview?
Step 5 Preview is cheaper: $1.43 against $3.00 per million tokens, blended at three input tokens per output token.
Which is faster, Qwen3.8 Max or Step 5 Preview?
Step 5 Preview streams faster: 87 against 37 output tokens per second.
Which has the larger context window?
Both accept 1M tokens of context.