BenchLeader
AlibabaAuto-detected

Qwen3.5-4B

Best configuration ranks #356 of 610 on the BenchLeader Index at 47.1 ±6.6. Released 24 Feb 2026.

Blended price
$0.060/M
$0.030 in · $0.150 out
Output speed
27 tok/s
First answer
74 s
Context
262k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index47
  2. Agents & tools49
  3. Maths31
  4. Knowledge36
  5. Instruction following53
  6. Multimodal50
  7. Long context58
  8. Composite46

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning43.9#42922 tok/s0.80 s$0.0001Agents & tools 43 · Composite 43 · Instruction following 37 · Knowledge 29 · Long context 43 · Maths 49 · Multimodal 46
defaultbest47.1#35627 tok/s74 s$0.0001Agents & tools 49 · Composite 46 · Instruction following 53 · Knowledge 36 · Long context 58 · Maths 31 · Multimodal 50

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoningdefaultSource
GPQA Diamond (AA)not in index71.2%#27877.1%#217Artificial Analysis
Humanity's Last Exam (AA)not in index8.0%#2849.9%#259Artificial Analysis

Agents & tools

Benchmarkno reasoningdefaultSource
Terminal-Bench Hard11.4%#20718.2%#166Artificial Analysis
τ²-Bench Telecom (AA)not in index87.7%#7692.1%#53Artificial Analysis

Maths

Benchmarkno reasoningdefaultSource
OTIS Mock AIME55.8%#156Epoch AI Benchmarking Hub
AIME 202689.7%#28MathArena
HMMT February 202672.9%#28MathArena

Knowledge

Benchmarkno reasoningdefaultSource
AA-Omniscience-74.7#448-58.4#377Artificial Analysis

Instruction following

Benchmarkno reasoningdefaultSource
IFBench33.3%#31552.0%#161Artificial Analysis

Multimodal

Benchmarkno reasoningdefaultSource
MMMU-Pro62.1%#16665.4%#142Artificial Analysis

Long context

Benchmarkno reasoningdefaultSource
AA-LCR34.7%#30563.0%#206Artificial Analysis

Composite

Benchmarkno reasoningdefaultSource
AA Intelligence Index10.8#30213.1#251Artificial Analysis

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.00011.4 min
Summarise a 30-page report12,000 / 600$0.00051.6 min
Code edit6,000 / 1,500$0.00042.1 min
Agentic coding session60,000 / 4,000$0.00243.7 min
Structured extraction2,000 / 200$0.00011.4 min

See also

Data as of 9 Sept 2026. Compare these configurations.