BenchLeader
AlibabaAuto-detected

qwen3-4b-instruct-2507

qwen3-4b-instruct-2507 is an Alibaba proprietary model, released 6 Aug 2025. Its best configuration ranks #501 of 372 on the BenchLeader Index at 41.4 ±3.8, in the lower half. It scores highest in maths (47) and lowest in long context (31). It has been measured at 2 reasoning-effort settings; tables show the best-scoring one. Last measured 6 Aug 2025.stale: no new result in six months

Blended price
Output speed
First answer
Context
262k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index41
  2. Reasoning36
  3. Agents & tools44
  4. Maths47
  5. Knowledge35
  6. Instruction following37
  7. Long context31
  8. Composite38

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
thinking40.2#525Agents & tools 35 · Composite 40 · Instruction following 51 · Knowledge 36 · Long context 45 · Maths 13
defaultbest41.4#501Agents & tools 44 · Composite 38 · Instruction following 37 · Knowledge 35 · Long context 31 · Maths 47 · Reasoning 36

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkthinkingdefaultSource
GPQA Diamond45.8%#224Epoch AI Benchmarking Hub
GPQA Diamond (AA)not in index66.7%#31951.7%#404Artificial Analysis
Humanity's Last Exam (AA)not in index6.2%#3394.5%#428Artificial Analysis

Agents & tools

BenchmarkthinkingdefaultSource
Terminal-Bench Hard1.5%#3284.5%#275Artificial Analysis
τ²-Bench Telecom (AA)not in index25.4%#28326.6%#272Artificial Analysis
BFCL Overall35.7%#41Berkeley Function Calling Leaderboard

Maths

BenchmarkthinkingdefaultSource
OTIS Mock AIME52.2%#167Epoch AI Benchmarking Hub
AIME 202682.5%#30MathArena
HMMT February 202653.0%#31MathArena
MathArena Apex2.1%#29MathArena

Knowledge

BenchmarkthinkingdefaultSource
AA-Omniscience-60.3#401-61.9#409Artificial Analysis

Instruction following

BenchmarkthinkingdefaultSource
IFBench49.8%#17433.5%#315Artificial Analysis

Long context

BenchmarkthinkingdefaultSource
AA-LCR37.3%#30111.3%#394Artificial Analysis

Composite

BenchmarkthinkingdefaultSource
AA Intelligence Index8.8#3566.7#437Artificial Analysis

See also

Data as of 10 Sept 2026. Compare these configurations.