BenchLeader
AlibabaOpen weightsReasoning model

Qwen3 235B A22B 2507

Qwen3 235B A22B 2507 is an Alibaba open-weights reasoning model, released 22 Jul 2025. Its best configuration (thinking reasoning effort) ranks #247 of 372 on the BenchLeader Index at 52.6 ±3.7, in the lower half. It scores highest in maths (62) and lowest in composite (45). At $0.747 per million tokens blended it is mid-priced. Output speed of 61 tokens per second puts it slower than most, with a first answer in 35.7 s. It has been measured at 2 reasoning-effort settings; tables show the best-scoring one. Last measured 2 Sept 2026.

Blended price
$0.747/M
$0.230 in · $2.30 out
Output speed
61 tok/s
measured by Artificial Analysis
First answer
36 s
first token 0.57 s
Context
256k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index53
  2. Reasoning59
  3. Coding50
  4. Agents & tools46
  5. Maths62
  6. Knowledge45
  7. Instruction following52
  8. Human preference59
  9. Long context60
  10. Composite45

Versions

Alibaba has shipped 2 models under this name. Each is ranked on its own results; a newer version often has fewer results so far, which holds its index nearer the average until more arrive.

ModelReleasedIndexRank
Qwen3 235B A22B 2507thinkingthis page22 Jul 202552.6#247
Qwen3 235B A22Bthinking28 Apr 202544.8#417

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
thinkingbest52.6#24761 tok/s36 s$0.0008Agents & tools 46 · Coding 50 · Composite 45 · Human preference 59 · Instruction following 52 · Knowledge 45 · Long context 60 · Maths 62 · Reasoning 59
default51.0#28158 tok/s2.38 s$0.0004Agents & tools 57 · Coding 52 · Composite 44 · Human preference 61 · Instruction following 48 · Knowledge 43 · Long context 43 · Reasoning 62

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkthinkingdefaultSource
GPQA Diamond80.0%#115Epoch AI Benchmarking Hub
LMArena Hard Prompts1417#1381448#94LMArena
GPQA Diamond (AA)not in index79.0%#20075.3%#243Artificial Analysis
Humanity's Last Exam (AA)not in index15.9%#19311.1%#240Artificial Analysis

Coding

BenchmarkthinkingdefaultSource
SciCode42.4%#99SciCode
WeirdML41.0%#9838.7%#111WeirdML
LMArena Coding1443#1351472#98LMArena
SciCode (AA)not in index41.4%#111Artificial Analysis

Agents & tools

BenchmarkthinkingdefaultSource
Terminal-Bench Hard13.6%#19515.2%#187Artificial Analysis
τ²-Bench Telecom (AA)not in index53.2%#18633.3%#238Artificial Analysis
BFCL Overall52.1%#21Berkeley Function Calling Leaderboard

Maths

BenchmarkthinkingdefaultSource
OTIS Mock AIME86.7%#73Epoch AI Benchmarking Hub

Knowledge

BenchmarkthinkingdefaultSource
SimpleQA Verified40.4%#40Epoch AI Benchmarking Hub
AA-Omniscience-46.5#297-44.0#284Artificial Analysis
MultiNRC27.1%#28Scale AI SEAL

Instruction following

BenchmarkthinkingdefaultSource
IFBench51.2%#16646.0%#195Artificial Analysis

Human preference

BenchmarkthinkingdefaultSource
LMArena Text1400#1331423#103LMArena

Long context

BenchmarkthinkingdefaultSource
Fiction.LiveBench 120k68.8%#8Fiction.live
AA-LCR72.0%#14433.9%#312Artificial Analysis

Composite

BenchmarkthinkingdefaultSource
AA Intelligence Index12.7#26412#276Artificial Analysis

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Google Vertex (US) (ZDR)40 tok/s0.58 s$0.220$0.880262k
Alibaba Cloud Int.39 tok/s0.51 s$0.149$0.598131k
Google Vertex (US) (ZDR)38 tok/s0.51 s$0.250$1.00262k
AtlasCloud31 tok/s1.01 s$0.200$0.880131kfp8
Nebius Token Factory31 tok/s0.70 s$0.200$0.600262kfp8
NovitaAI29 tok/s0.60 s$0.090$0.580131kfp8
StreamLake27 tok/s0.95 s$0.210$0.840128k
Parasail23 tok/s0.57 s$0.140$0.800131kfp8
Venice18 tok/s0.69 s$0.150$0.750128kfp8
DeepInfra12 tok/s0.32 s$0.090$0.550262kfp8
GMICloud10 tok/s1.30 s$0.087$0.350262kfp8

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$0.385$0.770$1.16$1.54Aug 26
input outputnow $0.087 in · $0.350 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.000840.6 s
Summarise a 30-page report12,000 / 600$0.004145.5 s
Code edit6,000 / 1,500$0.00481.0 min
Agentic coding session60,000 / 4,000$0.0231.7 min
Structured extraction2,000 / 200$0.000939.0 s

See also

Data as of 10 Sept 2026. Compare these configurations.