BenchLeader
AlibabaAuto-detected

Qwen2.5-Max

Not yet ranked: too few independent results so far.

Blended price
Output speed
First answer
Context
32k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Reasoning55
  2. Coding53
  3. Human preference55
  4. Composite39

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
maxbestCoding 53 · Composite 39 · Human preference 55 · Reasoning 55
default

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkmaxdefaultSource
LMArena Hard Prompts1385#162LMArena
GPQA Diamond (AA)not in index58.7%#366Artificial Analysis
Humanity's Last Exam (AA)not in index3.8%#477Artificial Analysis

Coding

BenchmarkmaxdefaultSource
LMArena Coding1402#171LMArena
SWE-bench Verified (any scaffold)not in index40.2%#51SWE-bench

Human preference

BenchmarkmaxdefaultSource
LMArena Text1374#157LMArena

Composite

See also

Data as of 9 Sept 2026. Compare these configurations.