BenchLeader

Cost to run a full benchmark suite

What Artificial Analysis actually paid to run its whole Intelligence Index on each configuration: a fixed batch of work, billed at list price. The spread between models is far wider than their per-token prices, because verbose reasoning multiplies the token count.

As of 21 Sept 2026, BenchLeader ranks Granite 4.2 3B first on the Cost to run a full benchmark suite board at $0.0060, ahead of GPT-5.6 Luna ($0.0098) and GPT-5.6 Luna ($0.010).

142 models · data as of 21 Sept 2026

Quality against what it costs to get there

Model size
40506070$0.0050$0.010$0.020$0.050$0.100$0.200$0.500$1.00$2.00$5.00$10.00What one full Intelligence Index run cost ($, log scale)BenchLeader IndexClaude Fable 5.1GPT-6 AstraClaude Fable 5Claude Opus 5Muse Spark 1.3GPT-5.6 SolGPT-5.5Kimi K3Qwen3 8Grok 4.6NEW Step 5 Preview
  • OpenAI31
  • Anthropic18
  • Alibaba15
  • Google12
  • Mistral AI8
  • SpaceXAI7
  • DeepSeek6
  • Meta5
  • Zhipu AI5
  • Moonshot AI4
  • NVIDIA4
  • IBM2
  • MiniMax2
  • Xiaomi2
  • Ant Group1
  • Arcee AI1
  • Celeris1
  • Meituan1
Dashed line: the frontier. Purple-ringed points are not beaten on both axes by any other model shown. Dashed ring: new in the last seven days.

The dashed staircase is the efficiency frontier: at each price, the best index anyone reaches. A model above and left of the rest gives you more quality per dollar of real work. Size bands and run costs come from Artificial Analysis, which is the only source that publishes what a fixed batch of work actually cost.

84 of 84
#
1Claude Fable 5thinkingAnthropic$8.7570.8$20.00
2Qwen3 8maxAlibaba$5.4166.5$2.67
3Claude Opus 4.8maxAnthropic$4.0864.1$10.00
4Claude Sonnet 4.6maxAnthropic$2.4957.5$6.00
5Claude Fable 5.1lowAnthropic$2.3767.4$20.00
6Qwen3.8 2.4T A95BAlibabaopen ↗$2.1665.1$3.00
7Quasar 438BMultiverse Computing$2.0257.1$0.900
8GLM 5.3maxZhipu AIopen ↗$2.0165.8$2.15
9Kimi K3maxMoonshot AIopen ↗$2.0067.1$6.00
10Gemini 3.5 FlashhighGoogle$1.5663.5$3.38
11Muse Spark 1.1xhighMeta$1.3861.7$2.00
12Muse Spark 1.3xhighMeta$1.3767.9$2.00
13Qwen3 7maxAlibaba$1.1560.9$3.75
14Claude Opus 5lowAnthropic$1.1063.3$10.00
15Nemotron 3 SuperthinkingNVIDIAopen ↗$1.0651.3$0.350
16Grok 4.5highSpaceXAI$1.0460.9$3.00
17Muse Spark 1.2xhighMeta$0.97564.0$2.00
18GLM 5.2maxZhipu AIopen ↗$0.96563.8$2.15
19Gemini 3.8 FlashmediumGoogle$0.93164.8$1.50
20Gemini 3.6 FlashhighGoogle$0.92961.7$1.50
21Gemini 3.7 FlashhighGoogle$0.92564.5$1.50
22GLM 5.1thinkingZhipu AIopen ↗$0.92257.7$2.15
23GPT-5.5mediumOpenAI$0.90264.5$11.25
24Qwen3.8 27BxhighAlibabaopen ↗$0.82059.1$1.06
25GPT-6 AstralowOpenAI$0.81868.2$20.00
26Kimi K2.6Moonshot AIopen ↗$0.80060.4$1.71
27NewStep 5 PreviewStepFun$0.71664.9$1.43
28GPT-5.5 InstantOpenAI$0.69159.0$11.25
29Gemini 3.1 ProGoogle$0.67563.8$4.50
30DeepSeek V4 PromaxDeepSeekopen ↗$0.67464.1$0.544
31Qwen3.6 27BthinkingAlibabaopen ↗$0.62154.8$1.35
32Nemotron 3 Ultra 550B A55BthinkingNVIDIAopen ↗$0.57957.5$1.00
33Claude Sonnet 4.5thinkingAnthropic$0.55753.3$6.00
34Qwen3 Coder NextAlibabaopen ↗$0.55242.5$0.450
35Kimi K2.7 CodeMoonshot AIopen ↗$0.54155.8$1.71
36Claude Sonnet 5lowAnthropic$0.509$4.00
37MiniMax M3MiniMaxopen ↗$0.50857.4$0.525
38Grok 4.6lowSpaceXAI$0.47560.9$3.00
39Qwen3.6 35B A3BthinkingAlibabaopen ↗$0.47553.0$0.557
40Qwen3.5 397B A17BthinkingAlibabaopen ↗$0.47556.0$0.387
41Mistral Medium 3.5Mistral AIopen$0.43751.2$3.00
42GPT-5.4 MinixhighOpenAI$0.41055.7$1.69
43Mistral Medium 3Mistral AI$0.39245.5$0.800
44Qwen3.8-Flash-NextAlibabaopen ↗$0.37261.5$0.230
45Trinity LargethinkingArcee AIopen$0.35347.6$0.413
46Qwen3.7 PlusAlibaba$0.32559.8$1.13
47Qwen3.5-122B-A10BthinkingAlibabaopen ↗$0.32553.4$1.10
48Deepseek v4 Flash VisionmaxDeepSeek$0.31459.7$0.262
49Ring-2.6-1TAnt Groupopen$0.28750.0$0.850
50DeepSeek V4.1 FlashmaxDeepSeekopen ↗$0.26561.8$0.262
51DeepSeek R1DeepSeekopen ↗$0.26248.7$0.315
52GPT-5.6 SollowOpenAI$0.26163.6$8.00
53GLM 5.3 FlashZhipu AIopen ↗$0.25363.9$0.238
54Gemini 2.5 ProGoogle$0.23254.4$3.44
55DeepSeek V4 FlashmaxDeepSeekopen ↗$0.22062.0$0.262
56Claude Haiku 4.5thinkingAnthropic$0.20847.4$2.00
57GPT-5.4 NanoxhighOpenAI$0.18354.5$0.463
58GPT-5.6 Terrano reasoningOpenAI$0.14050.7$4.50
59Ministral 3 14BMistral AIopen$0.13940.0$0.200
60Grok 4.3no reasoningSpaceXAI$0.13746.6$1.56

Cite as: BenchLeader, “Cost to run a full benchmark suite”, https://www.benchleader.com/leaderboards/cost-per-run, data as of 21 Sept 2026.