BenchLeader
AlibabaAuto-detected

Qwen3-1.7B

Best configuration ranks #557 of 610 on the BenchLeader Index at 37.8 ±3.5 (thinking reasoning effort). Released 29 Apr 2025.

Blended price
Output speed
First answer
Context
32k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index38
  2. Agents & tools34
  3. Knowledge27
  4. Instruction following31
  5. Long context25
  6. Composite36

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoningMaths 27 · Reasoning 25
thinkingbest37.8#557Agents & tools 34 · Composite 36 · Instruction following 31 · Knowledge 27 · Long context 25
default35.8#591Agents & tools 39 · Composite 35 · Instruction following 26 · Knowledge 24 · Long context 25 · Reasoning 30

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoningthinkingdefaultSource
GPQA Diamond30.6%#24938.0%#234Epoch AI Benchmarking Hub
GPQA Diamond (AA)not in index35.6%#46528.3%#499Artificial Analysis
Humanity's Last Exam (AA)not in index4.6%#4045.3%#358Artificial Analysis

Agents & tools

Benchmarkno reasoningthinkingdefaultSource
Terminal-Bench Hard0.0%#3450.0%#345Artificial Analysis
τ²-Bench Telecom (AA)not in index26.0%#27721.6%#305Artificial Analysis
BFCL Overall28.4%#53Berkeley Function Calling Leaderboard

Maths

Benchmarkno reasoningthinkingdefaultSource
OTIS Mock AIME8.1%#213Epoch AI Benchmarking Hub

Knowledge

Benchmarkno reasoningthinkingdefaultSource
AA-Omniscience-77.5#455-84.1#468Artificial Analysis

Instruction following

Benchmarkno reasoningthinkingdefaultSource
IFBench26.9%#36321.1%#389Artificial Analysis

Long context

Benchmarkno reasoningthinkingdefaultSource
AA-LCR0.0%#4190.0%#419Artificial Analysis

Composite

Benchmarkno reasoningthinkingdefaultSource
AA Intelligence Index5.2#5094.9#524Artificial Analysis

See also

Data as of 9 Sept 2026. Compare these configurations.