Token efficiency: quality per dollar of real work
BenchLeader Index divided by what it cost Artificial Analysis to run its whole Intelligence Index on that configuration. Unlike price per token this counts the tokens a model actually burns, so a model that thinks at length scores lower than its sticker price suggests.
As of 21 Sept 2026, BenchLeader ranks Granite 4.2 3B first on the Token efficiency: quality per dollar of real work board at 7625, ahead of GPT-5.6 Luna (5310) and GPT-5.6 Luna (4549).
128 models · data as of 21 Sept 2026
Quality against what it costs to get there
OpenAI31
Anthropic18
Alibaba15
Google12
Mistral AI8
SpaceXAI7DeepSeek6
Meta5
Zhipu AI5
Moonshot AI4
NVIDIA4
IBM2
MiniMax2
Xiaomi2
Ant Group1
Arcee AI1
- Celeris1
Meituan1
The dashed staircase is the efficiency frontier: at each price, the best index anyone reaches. A model above and left of the rest gives you more quality per dollar of real work. Size bands and run costs come from Artificial Analysis, which is the only source that publishes what a fixed batch of work actually cost.
- 1Granite 4.2 3B7625
- 2GPT-5.6 Luna (low)5310
- 3gpt-oss-20b (high)3945
- 4Nemotron 3 Nano 30B A3B (thinking)2812
- 5DeepSeek V32268
- 6Granite 4.2 8B1983
- 7Gemini 3.1 Flash Lite1383
- 8Mistral Small 3.11218
- 9MiMo-V2.5-Pro1093
- 10GPT-5 mini (high)1032
- 11Mistral Small 4 (thinking)1027
- 12Muse Glimmer (high)958
- 13Ministral 3 3B887
- 14LongCat 2.0853
- 15Celeris-1835
| # | |||||
|---|---|---|---|---|---|
| 1 | 7625 | 45.6 | $0.0060 | – | |
| 2 | 5310 | 52.2 | $0.0098 | – | |
| 3 | 3945 | 44.7 | $0.011 | – | |
| 4 | 2812 | 47.1 | $0.017 | – | |
| 5 | 2268 | 44.8 | $0.020 | – | |
| 6 | 1983 | 46.8 | $0.024 | – | |
| 7 | 1383 | 54.5 | $0.039 | – | |
| 8 | 1218 | 39.9 | $0.033 | – | |
| 9 | 1093 | 59.1 | $0.054 | – | |
| 10 | 1032 | 55.2 | $0.053 | – | |
| 11 | 1027 | 46.3 | $0.045 | – | |
| 12 | 958 | 52.8 | $0.055 | – | |
| 13 | 887 | 37.5 | $0.042 | – | |
| 14 | 853 | 50.3 | $0.059 | – | |
| 15 | Celeris-1Celeris | 835 | 41.9 | $0.050 | – |
| 16 | 634 | 51.1 | $0.081 | – | |
| 17 | 594 | 38.9 | $0.066 | – | |
| 18 | 576 | 44.1 | $0.077 | – | |
| 19 | 553 | 43.9 | $0.079 | – | |
| 20 | 551 | 56.0 | $0.102 | – | |
| 21 | 518 | 48.2 | $0.093 | – | |
| 22 | 462 | 49.6 | $0.107 | – | |
| 23 | 449 | 55.4 | $0.124 | – | |
| 24 | 441 | 46.0 | $0.104 | – | |
| 25 | 400 | 57.8 | $0.144 | – | |
| 26 | 352 | 58.1 | $0.165 | – | |
| 27 | 298 | 54.5 | $0.183 | – | |
| 28 | 287 | 40.0 | $0.139 | – | |
| 29 | 282 | 62.0 | $0.220 | – | |
| 30 | 252 | 63.9 | $0.253 | – | |
| 31 | 244 | 63.6 | $0.261 | – | |
| 32 | 235 | 54.4 | $0.232 | – | |
| 33 | 233 | 61.8 | $0.265 | – | |
| 34 | 228 | 47.4 | $0.208 | – | |
| 35 | 190 | 59.7 | $0.314 | – | |
| 36 | 186 | 48.7 | $0.262 | – | |
| 37 | 184 | 59.8 | $0.325 | – | |
| 38 | 174 | 50.0 | $0.287 | – | |
| 39 | 165 | 61.5 | $0.372 | – | |
| 40 | 164 | 53.4 | $0.325 | – | |
| 41 | 136 | 55.7 | $0.410 | – | |
| 42 | 135 | 47.6 | $0.353 | – | |
| 43 | 128 | 60.9 | $0.475 | – | |
| 44 | 118 | 56.0 | $0.475 | – | |
| 45 | 117 | 51.2 | $0.437 | – | |
| 46 | 116 | 45.5 | $0.392 | – | |
| 47 | 113 | 57.4 | $0.508 | – | |
| 48 | 112 | 53.0 | $0.475 | – | |
| 49 | 103 | 55.8 | $0.541 | – | |
| 50 | 99.3 | 57.5 | $0.579 | – | |
| 51 | 95.8 | 53.3 | $0.557 | – | |
| 52 | 95.1 | 64.1 | $0.674 | – | |
| 53 | 94.6 | 63.8 | $0.675 | – | |
| 54 | New | 90.7 | 64.9 | $0.716 | – |
| 55 | 88.2 | 54.8 | $0.621 | – | |
| 56 | 85.3 | 59.0 | $0.691 | – | |
| 57 | 83.4 | 68.2 | $0.818 | – | |
| 58 | 77.0 | 42.5 | $0.552 | – | |
| 59 | 75.5 | 60.4 | $0.800 | – | |
| 60 | 72.1 | 59.1 | $0.820 | – |
Cite as: BenchLeader, “Token efficiency: quality per dollar of real work”, https://www.benchleader.com/leaderboards/efficiency, data as of 21 Sept 2026.