Cost to run a full benchmark suite
What Artificial Analysis actually paid to run its whole Intelligence Index on each configuration: a fixed batch of work, billed at list price. The spread between models is far wider than their per-token prices, because verbose reasoning multiplies the token count.
As of 21 Sept 2026, BenchLeader ranks Granite 4.2 3B first on the Cost to run a full benchmark suite board at $0.0060, ahead of GPT-5.6 Luna ($0.0098) and GPT-5.6 Luna ($0.010).
142 models · data as of 21 Sept 2026
Quality against what it costs to get there
OpenAI31
Anthropic18
Alibaba15
Google12
Mistral AI8
SpaceXAI7DeepSeek6
Meta5
Zhipu AI5
Moonshot AI4
NVIDIA4
IBM2
MiniMax2
Xiaomi2
Ant Group1
Arcee AI1
- Celeris1
Meituan1
The dashed staircase is the efficiency frontier: at each price, the best index anyone reaches. A model above and left of the rest gives you more quality per dollar of real work. Size bands and run costs come from Artificial Analysis, which is the only source that publishes what a fixed batch of work actually cost.
- 1Granite 4.2 3B$0.0060
- 2GPT-5.6 Luna (low)$0.0098
- 3gpt-oss-20b (high)$0.011
- 4Nemotron 3 Nano 30B A3B (thinking)$0.017
- 5DeepSeek V3$0.020
- 6Granite 4.2 8B$0.024
- 7Mistral Small 3.1$0.033
- 8Gemini 3.1 Flash Lite$0.039
- 9Ministral 3 3B$0.042
- 10Mistral Small 4 (thinking)$0.045
- 11Celeris-1$0.050
- 12GPT-5 mini (high)$0.053
- 13MiMo-V2.5-Pro$0.054
- 14Muse Glimmer (high)$0.055
- 15LongCat 2.0$0.059
| # | ||||
|---|---|---|---|---|
| 1 | $8.75 | 70.8 | $20.00 | |
| 2 | $5.41 | 66.5 | $2.67 | |
| 3 | $4.08 | 64.1 | $10.00 | |
| 4 | $2.49 | 57.5 | $6.00 | |
| 5 | $2.37 | 67.4 | $20.00 | |
| 6 | $2.16 | 65.1 | $3.00 | |
| 7 | $2.02 | 57.1 | $0.900 | |
| 8 | $2.01 | 65.8 | $2.15 | |
| 9 | $2.00 | 67.1 | $6.00 | |
| 10 | $1.56 | 63.5 | $3.38 | |
| 11 | $1.38 | 61.7 | $2.00 | |
| 12 | $1.37 | 67.9 | $2.00 | |
| 13 | $1.15 | 60.9 | $3.75 | |
| 14 | $1.10 | 63.3 | $10.00 | |
| 15 | $1.06 | 51.3 | $0.350 | |
| 16 | $1.04 | 60.9 | $3.00 | |
| 17 | $0.975 | 64.0 | $2.00 | |
| 18 | $0.965 | 63.8 | $2.15 | |
| 19 | $0.931 | 64.8 | $1.50 | |
| 20 | $0.929 | 61.7 | $1.50 | |
| 21 | $0.925 | 64.5 | $1.50 | |
| 22 | $0.922 | 57.7 | $2.15 | |
| 23 | $0.902 | 64.5 | $11.25 | |
| 24 | $0.820 | 59.1 | $1.06 | |
| 25 | $0.818 | 68.2 | $20.00 | |
| 26 | $0.800 | 60.4 | $1.71 | |
| 27 | New | $0.716 | 64.9 | $1.43 |
| 28 | $0.691 | 59.0 | $11.25 | |
| 29 | $0.675 | 63.8 | $4.50 | |
| 30 | $0.674 | 64.1 | $0.544 | |
| 31 | $0.621 | 54.8 | $1.35 | |
| 32 | $0.579 | 57.5 | $1.00 | |
| 33 | $0.557 | 53.3 | $6.00 | |
| 34 | $0.552 | 42.5 | $0.450 | |
| 35 | $0.541 | 55.8 | $1.71 | |
| 36 | $0.509 | – | $4.00 | |
| 37 | $0.508 | 57.4 | $0.525 | |
| 38 | $0.475 | 60.9 | $3.00 | |
| 39 | $0.475 | 53.0 | $0.557 | |
| 40 | $0.475 | 56.0 | $0.387 | |
| 41 | $0.437 | 51.2 | $3.00 | |
| 42 | $0.410 | 55.7 | $1.69 | |
| 43 | $0.392 | 45.5 | $0.800 | |
| 44 | $0.372 | 61.5 | $0.230 | |
| 45 | $0.353 | 47.6 | $0.413 | |
| 46 | $0.325 | 59.8 | $1.13 | |
| 47 | $0.325 | 53.4 | $1.10 | |
| 48 | $0.314 | 59.7 | $0.262 | |
| 49 | $0.287 | 50.0 | $0.850 | |
| 50 | $0.265 | 61.8 | $0.262 | |
| 51 | $0.262 | 48.7 | $0.315 | |
| 52 | $0.261 | 63.6 | $8.00 | |
| 53 | $0.253 | 63.9 | $0.238 | |
| 54 | $0.232 | 54.4 | $3.44 | |
| 55 | $0.220 | 62.0 | $0.262 | |
| 56 | $0.208 | 47.4 | $2.00 | |
| 57 | $0.183 | 54.5 | $0.463 | |
| 58 | $0.140 | 50.7 | $4.50 | |
| 59 | $0.139 | 40.0 | $0.200 | |
| 60 | $0.137 | 46.6 | $1.56 |
Cite as: BenchLeader, “Cost to run a full benchmark suite”, https://www.benchleader.com/leaderboards/cost-per-run, data as of 21 Sept 2026.