BenchLeader
AlibabaOpen weightsReasoning model

Qwen3 30B A3B 2507

Qwen3 30B A3B 2507 is an Alibaba open-weights reasoning model, released 30 Jul 2025. Its best configuration ranks #377 of 372 on the BenchLeader Index at 46.3 ±5.3, in the lower half. It scores highest in coding (58) and lowest in knowledge (33). At $0.150 per million tokens blended it is among the cheapest fifth of ranked models. Output speed of 141 tokens per second puts it in the fastest quarter, with a first answer in 1.9 s. It has been measured at 2 reasoning-effort settings; tables show the best-scoring one. Last measured 2 Sept 2026.

Blended price
$0.150/M
$0.100 in · $0.300 out
Output speed
141 tok/s
measured by Artificial Analysis
First answer
1.88 s
first token 0.48 s
Context
262k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index46
  2. Reasoning47
  3. Coding58
  4. Agents & tools48
  5. Maths52
  6. Knowledge33
  7. Instruction following37
  8. Human preference56
  9. Long context39
  10. Composite39

Versions

Alibaba has shipped 2 models under this name. Each is ranked on its own results; a newer version often has fewer results so far, which holds its index nearer the average until more arrive.

ModelReleasedIndexRank
Qwen3 30B A3B 2507this page30 Jul 202546.3#377
Qwen3 30B A3B42.4#471

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
thinking45.0#409149 tok/s16 s$0.0008Agents & tools 38 · Coding 35 · Composite 41 · Instruction following 52 · Knowledge 37 · Long context 57 · Maths 38 · Reasoning 52
defaultbest46.3#377141 tok/s1.88 s$0.0001Agents & tools 48 · Coding 58 · Composite 39 · Human preference 56 · Instruction following 37 · Knowledge 33 · Long context 39 · Maths 52 · Reasoning 47

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkthinkingdefaultSource
GPQA Diamond70.1%#16055.6%#192Epoch AI Benchmarking Hub
LMArena Hard Prompts1407#147LMArena
GPQA Diamond (AA)not in index70.7%#28565.9%#327Artificial Analysis
Humanity's Last Exam (AA)not in index10.3%#2586.9%#312Artificial Analysis
Kagi LLM Benchmark54.9%#67Kagi LLM Benchmark

Coding

BenchmarkthinkingdefaultSource
SciCode33.3%#140SciCode
LMArena Coding1439#139LMArena
SciCode (AA)not in index33.0%#141Artificial Analysis

Agents & tools

BenchmarkthinkingdefaultSource
Terminal-Bench Hard5.3%#2686.1%#259Artificial Analysis
τ²-Bench Telecom (AA)not in index28.1%#26110.2%#375Artificial Analysis
BFCL Overall41.4%#33Berkeley Function Calling Leaderboard

Maths

BenchmarkthinkingdefaultSource
OTIS Mock AIME70.3%#12662.2%#146Epoch AI Benchmarking Hub
AIME 202688.3%#29MathArena
HMMT February 202678.8%#27MathArena
MathArena Apex0.5%#41MathArena

Knowledge

BenchmarkthinkingdefaultSource
AA-Omniscience-56.1#373-66.2#435Artificial Analysis

Instruction following

BenchmarkthinkingdefaultSource
IFBench50.7%#16933.1%#323Artificial Analysis

Human preference

BenchmarkthinkingdefaultSource
LMArena Text1382#152LMArena

Long context

BenchmarkthinkingdefaultSource
AA-LCR61.3%#21426.3%#339Artificial Analysis

Composite

BenchmarkthinkingdefaultSource
AA Intelligence Index9.8#3247.5#400Artificial Analysis

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Alibaba Cloud Int.69 tok/s0.32 s$0.130$0.520131k
StreamLake51 tok/s0.70 s$0.048$0.193128k
Nebius Token Factory23 tok/s0.48 s$0.100$0.300262kfp8
SiliconFlow16 tok/s1.46 s$0.090$0.300262kfp8

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$0.053$0.106$0.159$0.212Jun 26Jul 26Aug 26Sept 26
input outputnow $0.048 in · $0.193 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.010 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0001$0.00014.0 s
Summarise a 30-page report12,000 / 600$0.0014$0.00066.1 s
Code edit6,000 / 1,500$0.0011$0.000612.5 s
Agentic coding session60,000 / 4,000$0.0072$0.003230.3 s
Structured extraction2,000 / 200$0.0003$0.00013.3 s

See also

Data as of 10 Sept 2026. Compare these configurations.