BenchLeader
AlibabaOpen weights: Qwen/Qwen3.8-27BReasoning modelAuto-detected

Qwen 3.8 27B

Best configuration ranks #148 of 610 on the BenchLeader Index at 57.1 ±4.9. Last measured 8 Sept 2026. Released 14 Aug 2026.

Blended price
$1.13/M
$0.500 in · $3.00 out
Output speed
46 tok/s
First answer
47 s
first token 1.48 s
Context
262k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index57
  2. Reasoning63
  3. Coding66
  4. Agents & tools48
  5. Maths40
  6. Knowledge59
  7. Human preference63
  8. Multimodal62
  9. Long context68
  10. Composite62

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning52.6#24554 tok/s3.87 s$0.0011Coding 38 · Composite 57 · Knowledge 60 · Long context 61 · Multimodal 54
low54.3#20352 tok/s42 s$0.0011Coding 44 · Composite 66 · Knowledge 51 · Long context 65 · Multimodal 58
medium53.9#21450 tok/s44 s$0.0011Coding 41 · Composite 68 · Knowledge 47 · Long context 66 · Multimodal 59
xhigh52.0#26046 tok/s47 s$0.0015Agents & tools 47 · Coding 57 · Knowledge 55
defaultbest57.1#14846 tok/s47 s$0.0011Agents & tools 48 · Coding 66 · Composite 62 · Human preference 63 · Knowledge 59 · Long context 68 · Maths 40 · Multimodal 62 · Reasoning 63

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoninglowmediumxhighdefaultSourceTrend
LMArena Hard Prompts1460#79LMArena
LiveBench Reasoningnot in index80.0%#39LiveBench
GPQA Diamond (AA)not in index81.8%#17284.5%#13684.5%#13690.5%#59Artificial Analysis
Humanity's Last Exam (AA)not in index12.1%#21814.0%#20114.1%#19933.9%#85Artificial Analysis
GPQA Diamond (Vals)not in index88.9%#31Vals AI

Coding

Benchmarkno reasoninglowmediumxhighdefaultSourceTrend
SciCode35.6%#13139.8%#11138.1%#12144.7%#87SciCode
LMArena Coding1494#72LMArena
LMArena WebDev1592#15LMArena
LiveBench Codingnot in index75.7%#36LiveBench
SciCode (AA)not in index36.2%#13340.0%#11639.0%#11946.6%#85Artificial Analysis
LiveCodeBench84.0%#41Vals AI
SWE-bench (Vals)not in index86.0%#15Vals AI

Agents & tools

Benchmarkno reasoninglowmediumxhighdefaultSourceTrend
LMArena Agent0.1#25LMArena
LiveBench Agentic Codingnot in index61.4%#9LiveBench
Terminal-Bench 2.1 (Vals)58.4%#32Vals AI

Maths

Benchmarkno reasoninglowmediumxhighdefaultSourceTrend
ProofBench16.0%#41Vals AI
LiveBench Mathematicsnot in index86.2%#38LiveBench

Knowledge

Benchmarkno reasoninglowmediumxhighdefaultSourceTrend
LiveBench Data Analysisnot in index76.6%#22LiveBench
AA-Omniscience-8.0#127-26.7#204-36.1#239-10.0#139Artificial Analysis
MMLU-Pro84.3%#62Vals AI
LegalBench82.4%#64Vals AI
TaxEval70.8%#83Vals AI

Instruction following

Benchmarkno reasoninglowmediumxhighdefaultSourceTrend
LiveBench Languagenot in index74.3%#41LiveBench

Human preference

Benchmarkno reasoninglowmediumxhighdefaultSourceTrend
LMArena Text1436#84LMArena

Multimodal

Benchmarkno reasoninglowmediumxhighdefaultSourceTrend
LMArena Vision1279#28LMArena
MMMU-Pro69.9%#11973.8%#9374.2%#8776.3%#64Artificial Analysis

Long context

Benchmarkno reasoninglowmediumxhighdefaultSourceTrend
AA-LCR69.3%#17077.3%#9179.7%#5982.0%#25Artificial Analysis

Composite

Benchmarkno reasoninglowmediumxhighdefaultSourceTrend
LiveBench75.3%#28LiveBench
AA Intelligence Index22.4#14629.2#8430.7#7633.9#62Artificial Analysis
Vals Indexnot in index48.5#29Vals AI

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Venice87 tok/s1.49 s$0.450$3.20262kfp8
AkashML64 tok/s0.74 s$0.250$2.20262kfp8
Phala53 tok/s1.08 s$0.300$3.00262k
CoreWeave53 tok/s0.42 s$0.400$3.00262kfp8
Ionstream51 tok/s0.73 s$0.350$2.55262kfp8
Reka AI43 tok/s0.69 s$0.214$2.55262kfp8
Parasail43 tok/s1.48 s$0.240$2.20262kfp8
Alibaba Cloud Int.43 tok/s1.45 s$0.425$2.551M
Chutes30 tok/s2.17 s$0.320$2.50262kfp8
NovitaAI25 tok/s1.53 s$0.420$3.001M
Cloudflare25 tok/s1.62 s$0.450$3.20262k
io.net22 tok/s2.42 s$0.300$2.8066kfp8
Darkbloom16 tok/s1.74 s$0.150$2.00262kfp4

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$0.880$1.76$2.64$3.52Aug 26
input outputnow $0.350 in · $2.55 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.040 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0011$0.001053.9 s
Summarise a 30-page report12,000 / 600$0.0078$0.00371.0 min
Code edit6,000 / 1,500$0.0075$0.00541.3 min
Agentic coding session60,000 / 4,000$0.042$0.0212.2 min
Structured extraction2,000 / 200$0.0016$0.000951.8 s

See also

Data as of 9 Sept 2026. Compare these configurations.