BenchLeader

Qwen3.8 Max

Reasoning effort

Qwen3.8 Max is an Alibaba proprietary reasoning model, released 3 Aug 2026. Its best configuration (max reasoning effort) ranks #52 of 760 on the BenchLeader Index at 64.2 ±3.4, in the upper half. The ± is the point: 109 other configurations score within that range, so they and this one cannot be told apart on quality alone — price and speed are what separate them. It scores highest in reasoning (72) and lowest in agents & tools (58). At $3.00 per million tokens blended it is pricier than most ranked models. Output speed of 37 tokens per second puts it in the slowest quarter, with a first answer in 58.9 s. It has been measured at 3 reasoning-effort settings; this summary describes the best-scoring one, and the tabs above switch between them. Last measured 9 Oct 2026.

Blended price
$3.00/M
$2.00 in · $6.00 out
Output speed
37 tok/s
measured by Artificial Analysis
First answer
59 s
first token 2.70 s
Context
1M
Full answer
73 s
median, reasoning included
Cost per run
$5.41
one full Intelligence Index run
Released
3 Aug 2026
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index64
  2. Reasoning72
  3. Coding65
  4. Agents & tools58
  5. Maths65
  6. Knowledge63
  7. Human preference68
  8. Multimodal66
  9. Long context65
  10. Composite71

Versions

Alibaba has shipped 2 models under this name. Each is ranked on its own results; a newer version often has fewer results so far, which holds its index nearer the average until more arrive.

ModelReleasedIndexRank
Qwen3.8 Maxmaxthis page3 Aug 202664.2#52
Qwen3.6 Maxmax20 Apr 202662.5#77

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Not stated” collects results from publishers that did not say which setting they used; for a reasoning model that is usually its thinking mode, but we do not assume it. Pick a setting here or at the top of the page to see its price, speed and category scores.

EffortIndexRankSpeedFirst answerChat reply costCategories
xhigh57.3#17035 tok/s2.70 s$0.0026Knowledge 55 · Maths 63 · Reasoning 67
maxbest64.2#5237 tok/s59 s$0.0026Agents & tools 58 · Coding 65 · Composite 71 · Human preference 68 · Knowledge 63 · Long context 65 · Maths 65 · Multimodal 66 · Reasoning 72
not stated––35 tok/s2.70 s$0.0026Coding 67

How its index has moved

27 Sept 2026 to 8 Oct 2026
546269
  • max
  • xhigh

The index is recomputed from scratch every day, so a line moves when a new benchmark result lands, when a publisher revises a score, or when the models it is normalised against change. Early movement usually means the score is still settling.

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkxhighmaxnot statedSourceTrend
GPQA Diamond92.7%#24––Epoch AI Benchmarking Hub
LMArena Hard Prompts–1504#25–LMArena
LiveBench Reasoningnot in index–88.2%#22–LiveBench
GPQA Diamond (AA)not in index–92.8%#28–Artificial Analysis
Humanity's Last Exam (AA)not in index–43.1%#55–Artificial Analysis
GPQA Diamond (Vals)not in index–93.7%#6–Vals AI
CritPt–20.0%#57–Artificial Analysis
MysteryMechanismnot in index–23.9%#14–Vals AI
Chess Puzzlesnot in index40.0%#24––Epoch AI Benchmarking Hub
Mystery Game Puzzlesnot in index38.0%#15––Epoch AI Benchmarking Hub
LMCAnot in index46.2%#69––Epoch AI Benchmarking Hub
DTBenchnot in index92.0%#56––Epoch AI Benchmarking Hub

Coding

Benchmarkxhighmaxnot statedSourceTrend
SciCode–53.2%#7352.1%#79SciCode
LMArena Coding–1524#29–LMArena
LMArena WebDev–1672#111674#10LMArena
LiveBench Codingnot in index–72.9%#54–LiveBench
SciCode (AA)not in index–53.2%#73–Artificial Analysis
LiveCodeBench–87.8%#10–Vals AI
IOI–68.9%#11–Vals AI
SWE-bench (Vals)not in index–85.6%#17–Vals AI
Code Migrationnot in index–24.0%#41–Vals AI
ProgramBenchnot in index–0.0%#23–Vals AI
Vibe Code Bench 1-100not in index–12.8%#17–Vals AI
Vibe Code Bench v1.1not in index–64.7%#38–Vals AI
DeepSWE v1.1not in index57.5%#36––Epoch AI Benchmarking Hub
FrontierSWEnot in index17.8%#16––Epoch AI Benchmarking Hub
Terminal-Bench 4.0 (AA)not in index–38.9%#30–Artificial Analysis
Terminal-Bench 2.1 (AA)not in index–88.8%#9–Artificial Analysis

Agents & tools

Benchmarkxhighmaxnot statedSourceTrend
Terminal-Bench–27.0%#63–Terminal-Bench
APEX-Agents–63.3%#11–Mercor
LMArena Agent–2.3#22–LMArena
LiveBench Agentic Codingnot in index–64.7%#7–LiveBench
Terminal-Bench 2.1 (Vals)–67.4%#35–Vals AI
GDPval-AA v2.1not in index–58.6%#14–Artificial Analysis
τ³-Banking (AA)–51.3%#1–Artificial Analysis
ITBench SRE (AA)not in index–40.3%#19–Artificial Analysis
Analyst Agent (AA)not in index–45.0%#12–Artificial Analysis
APEX-Agents (AA)not in index–42.4%#2–Artificial Analysis
CyberBenchnot in index–28.6%#41–Vals AI
Finance Agent v2not in index–50.6%#35–Vals AI
Harvey's Legal Agent Benchmarknot in index–10.4%#14–Vals AI
Legal Research Benchnot in index–47.6%#11–Vals AI
SkillsBenchnot in index–42.0%#27–Vals AI
Tax Agent Benchnot in index–66.0%#15–Vals AI
Terminal-Bench 4.0 (Vals)not in index–34.3%#12–Vals AI
Terminal-Bench Sciencenot in index–1.4%#27–Vals AI
AutomationBenchnot in index–56.2%#30–Artificial Analysis
GDP.pdfnot in index–22.8%#18–Artificial Analysis
EnterpriseOps-Gymnot in index–47.6%#6–Artificial Analysis
AA-Briefcase v1.1not in index–1617#11–Artificial Analysis

Maths

Benchmarkxhighmaxnot statedSourceTrend
FrontierMath Tiers 1–374.7%#19––Epoch AI Benchmarking Hub
FrontierMath Tier 446.3%#25––Epoch AI Benchmarking Hub
OTIS Mock AIME100.0%#1––Epoch AI Benchmarking Hub
ProofBench–58.0%#19–Vals AI
LiveBench Mathematicsnot in index–91.3%#29–LiveBench
LMArena Maths–1497#19–LMArena

Knowledge

Benchmarkxhighmaxnot statedSourceTrend
SimpleQA Verified47.3%#32––Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index–78.4%#23–LiveBench
AA-Omniscience–12.0#86–Artificial Analysis
MMLU-Pro–88.6%#18–Vals AI
LegalBench–83.6%#48–Vals AI
CorpFin–65.8%#26–Vals AI
TaxEval–75.5%#17–Vals AI
Excel Modeling Benchmarknot in index–60.1%#32–Vals AI
MedCodenot in index–40.7%#57–Vals AI
MedScribenot in index–85.0%#30–Vals AI
Public Benefits Benchnot in index–67.1%#13–Vals AI
GDP.pdfnot in index23.2%#19––Epoch AI Benchmarking Hub
AA-Omniscience: accuracynot in index–31.9%#161–Artificial Analysis
AA-Omniscience: non-hallucinationnot in index–71.2%#34–Artificial Analysis

Instruction following

Benchmarkxhighmaxnot statedSourceTrend
LiveBench Languagenot in index–79.7%#33–LiveBench
LiveBench Instruction Followingnot in index–74.1%#13–LiveBench
LMArena Instruction Followingnot in index–1474#28–LMArena

Human preference

Benchmarkxhighmaxnot statedSourceTrend
LMArena Text–1483#22–LMArena
LMArena Creative Writingnot in index–1470#16–LMArena
LMArena Multi-turnnot in index–1492#19–LMArena
LMArena Longer Queriesnot in index–1492#25–LMArena

Multimodal

Benchmarkxhighmaxnot statedSourceTrend
LMArena Vision–1314#9–LMArena
MMMU-Pro–82.8%#29–Artificial Analysis
MMMU-Pro (Vals)not in index–88.0%#12–Vals AI
MortgageTaxnot in index–64.0%#48–Vals AI
SAGEnot in index–51.3%#14–Vals AI

Long context

Benchmarkxhighmaxnot statedSourceTrend
AA-LCR–80.3%#66–Artificial Analysis
MLCRnot in index–20.0%#19–Artificial Analysis

Composite

Benchmarkxhighmaxnot statedSourceTrend
Epoch Capabilities Indexnot in index–156.4#22–Epoch AI Benchmarking Hub
LiveBench–78.5%#16–LiveBench
AA Intelligence Index v4.3.2–45.4#33–Artificial Analysis
Vals Indexnot in index–48.3#28–Vals AI
Vals Multimodal Indexnot in index–65.4%#12–Vals AI

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Alibaba Cloud Int.38 tok/s1.78 s$2.00$6.001M–
Alibaba Cloud Int.34 tok/s2.70 s$2.00$6.001M–

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.250 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0026$0.00211.1 min
Summarise a 30-page report12,000 / 600$0.028$0.0121.3 min
Code edit6,000 / 1,500$0.021$0.0131.7 min
Agentic coding session60,000 / 4,000$0.144$0.0652.8 min
Structured extraction2,000 / 200$0.0052$0.00261.1 min

See also

Data as of 11 Oct 2026. Compare these configurations.

Cite as: BenchLeader, “Qwen3.8 Max: benchmarks, pricing, speed and rank”, https://www.benchleader.com/models/qwen3-8-max, data as of 11 Oct 2026.