BenchLeader

Qwen3 235B A22B

Best configuration ranks #267 of 610 on the BenchLeader Index at 51.7 ±3.6 (thinking reasoning effort). Last measured 2 Sept 2026. Released 28 Apr 2025.

Blended price
$0.975/M
$0.300 in · $3.00 out
Output speed
60 tok/s
First answer
37 s
first token 2.37 s
Context
256k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index52
  2. Reasoning60
  3. Coding50
  4. Agents & tools45
  5. Maths52
  6. Knowledge46
  7. Instruction following52
  8. Human preference58
  9. Long context60
  10. Composite45

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning58 tok/s2.37 s$0.0005Reasoning 51
thinkingbest51.7#26760 tok/s37 s$0.0010Agents & tools 45 · Coding 50 · Composite 45 · Human preference 58 · Instruction following 52 · Knowledge 46 · Long context 60 · Maths 52 · Reasoning 60
default48.7#31762 tok/s2.67 s$0.0005Agents & tools 57 · Coding 48 · Composite 44 · Human preference 61 · Instruction following 38 · Knowledge 48 · Long context 42 · Maths 57 · Reasoning 42

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoningthinkingdefaultSource
GPQA Diamond80.0%#11370.7%#156Epoch AI Benchmarking Hub
SimpleBench31.0%#66SimpleBench
LMArena Hard Prompts1417#1361448#94LMArena
GPQA Diamond (AA)not in index79.0%#19775.3%#240Artificial Analysis
Humanity's Last Exam (AA)not in index15.9%#18911.1%#235Artificial Analysis
GPQA Diamond (Vals)not in index70.2%#89Vals AI
Kagi LLM Benchmark55.0%#6669.4%#29Kagi LLM Benchmark
ARC-AGI-111.0%#168ARC Prize
ARC-AGI-21.3%#153ARC Prize

Coding

Benchmarkno reasoningthinkingdefaultSource
SciCode42.4%#97SciCode
WeirdML41.0%#9638.7%#109WeirdML
LMArena Coding1443#1341472#100LMArena
SciCode (AA)not in index41.4%#109Artificial Analysis
LiveCodeBench70.6%#84Vals AI
IOI0.0%#58Vals AI
SWE-Bench Pro21.4%#18Scale AI SEAL
Aider Polyglot59.6%#16Aider polyglot leaderboard

Agents & tools

Benchmarkno reasoningthinkingdefaultSource
Terminal-Bench Hard13.6%#19415.2%#186Artificial Analysis
τ²-Bench Telecom (AA)not in index53.2%#18433.3%#237Artificial Analysis
BFCL Overall52.1%#21Berkeley Function Calling Leaderboard

Maths

Benchmarkno reasoningthinkingdefaultSource
OTIS Mock AIME86.7%#71Epoch AI Benchmarking Hub
MATH Level 568.9%#39Epoch AI Benchmarking Hub
AIME (Vals)84.0%#40Vals AI
MGSM92.5%#17Vals AI
MathArena Apex5.2%#25MathArena

Knowledge

Benchmarkno reasoningthinkingdefaultSource
SimpleQA Verified40.4%#38Epoch AI Benchmarking Hub
AA-Omniscience-44.7#283-44.0#279Artificial Analysis
MMLU-Pro81.3%#79Vals AI
LegalBench80.2%#82Vals AI
TaxEval70.7%#86Vals AI
MedQA90.6%#44Vals AI
MultiNRC27.1%#2817.6%#36Scale AI SEAL

Instruction following

Benchmarkno reasoningthinkingdefaultSource
IFBench51.2%#16546.0%#194Artificial Analysis
MultiChallenge41.2%#26Scale AI SEAL

Human preference

Benchmarkno reasoningthinkingdefaultSource
LMArena Text1400#1321423#104LMArena

Long context

Benchmarkno reasoningthinkingdefaultSource
Fiction.LiveBench 120k68.8%#8Fiction.live
AA-LCR72.0%#14033.9%#308Artificial Analysis

Composite

Benchmarkno reasoningthinkingdefaultSource
Epoch Capabilities Indexnot in index143.9#78139.4#103Epoch AI Benchmarking Hub
AA Intelligence Index12.7#25812#272Artificial Analysis

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Alibaba Cloud Int.31 tok/s0.50 s$0.455$1.82131k

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.001041.8 s
Summarise a 30-page report12,000 / 600$0.005446.8 s
Code edit6,000 / 1,500$0.00631.0 min
Agentic coding session60,000 / 4,000$0.0301.7 min
Structured extraction2,000 / 200$0.001240.2 s

See also

Data as of 9 Sept 2026. Compare these configurations.