BenchLeader
AlibabaAuto-detected

Qwen3

Qwen3 is an Alibaba proprietary model. Its best configuration (max reasoning effort) ranks #192 of 372 on the BenchLeader Index at 55.0 ±6.8, in the lower half. It scores highest in instruction following (69) and lowest in knowledge (47). At $2.40 per million tokens blended it is pricier than most ranked models. Output speed of 55 tokens per second puts it slower than most, with a first answer in 40.1 s.

Blended price
$2.40/M
$1.20 in · $6.00 out
Output speed
55 tok/s
measured by Artificial Analysis
First answer
40 s
first token 1.02 s
Context
262k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index55
  2. Agents & tools55
  3. Knowledge47
  4. Instruction following69
  5. Long context64
  6. Composite56

Versions

Alibaba has shipped 5 models under this name. Each is ranked on its own results; a newer version often has fewer results so far, which holds its index nearer the average until more arrive.

ModelReleasedIndexRank
Qwen3 8max2 Aug 202663.6#44
Qwen3 7max19 May 202661.3#73
Qwen3 6max20 Apr 202663.3#49
Qwen3maxthis page55.0#192
Qwen3 5max

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkmaxSource
GPQA Diamond (AA)not in index86.1%#114Artificial Analysis
Humanity's Last Exam (AA)not in index28.0%#122Artificial Analysis

Agents & tools

Knowledge

BenchmarkmaxSource
AA-Omniscience-36.6#247Artificial Analysis

Instruction following

BenchmarkmaxSource
IFBench70.8%#64Artificial Analysis

Long context

BenchmarkmaxSource
AA-LCR74.3%#119Artificial Analysis

Composite

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Alibaba Cloud Int.30 tok/s1.02 s$0.780$3.90262k

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$1.65$3.30$4.95$6.60Feb 26Mar 26Apr 26May 26Jun 26Jul 26Aug 26Sept 26
input outputnow $0.780 in · $3.90 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.002345.5 s
Summarise a 30-page report12,000 / 600$0.01851.0 s
Code edit6,000 / 1,500$0.0161.1 min
Agentic coding session60,000 / 4,000$0.0961.9 min
Structured extraction2,000 / 200$0.003643.7 s

See also

Data as of 10 Sept 2026. Compare with another model.