BenchLeader

Step3 VL 10B

Best configuration ranks #447 of 610 on the BenchLeader Index at 43.0 ±6.9.

Blended price
Output speed
First answer
Context
66k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index43
  2. Agents & tools38
  3. Knowledge36
  4. Instruction following51
  5. Multimodal48
  6. Long context25
  7. Composite39

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkdefaultSource
GPQA Diamond (AA)not in index69.0%#295Artificial Analysis
Humanity's Last Exam (AA)not in index10.8%#244Artificial Analysis

Agents & tools

Knowledge

BenchmarkdefaultSource
AA-Omniscience-59#381Artificial Analysis

Instruction following

BenchmarkdefaultSource
IFBench50.2%#171Artificial Analysis

Multimodal

BenchmarkdefaultSource
MMMU-Pro64.0%#154Artificial Analysis

Long context

BenchmarkdefaultSource
AA-LCR0.0%#419Artificial Analysis

Composite

BenchmarkdefaultSource
AA Intelligence Index7.7#384Artificial Analysis

See also

Data as of 9 Sept 2026. Compare with another model.