What more thinking costs
Most frontier models let you choose how hard to think, and the setting is usually presented as a quality dial. It is also a price dial, and the two do not move together. Every figure here is what it cost Artificial Analysis to run its whole Intelligence Index at that setting — a fixed batch of work, billed at list price — so a model that reasons at length pays for it in a way its per-token price never shows.
The spread is wide. GLM 5.3 buys 6.6 index points for every doubling of spend; Qwen3.8 27B buys -7.1. Turning Qwen3.8 27B all the way up costs 2× its cheapest setting and moves it backwards by 9.2 points.
18 models whose settings span at least a doubling in cost · data as of 11 Oct 2026
GLM 5.3
6.6 points per doubling · 2.4× for +8.1
GPT-5.6 Terra
4.2 points per doubling · 10× for +14.0
Claude Fable 5.1
3.3 points per doubling · 3.2× for +5.5
Claude Opus 5
2.7 points per doubling · 5.3× for +6.5
Claude Sonnet 5.5
2.3 points per doubling · 16× for +9.2
Claude Opus 5.5
2.1 points per doubling · 11× for +7.1
GPT-6 Astra
2.0 points per doubling · 4.0× for +4.0
GPT-6 Sol
1.9 points per doubling · 7.9× for +5.7
GPT-5.5
1.9 points per doubling · 2.9× for +2.9
GPT-5.6 Luna
1.8 points per doubling · 18× for +7.7
GPT-5.6 Sol
1.8 points per doubling · 7.6× for +5.2
GPT-6.1 Sol
1.7 points per doubling · 5.5× for +4.1
Claude Sonnet 5
1.6 points per doubling · 10× for +5.3
Grok 4.6
1.4 points per doubling · 4.7× for +3.2
GPT-6 Luna
1.3 points per doubling · 15× for +5.2
Grok 4.7
-0.8 points per doubling · 3.0× for -1.3
Claude Haiku 5.5
-2.4 points per doubling · 2.7× for -3.4
Qwen3.8 27B
-7.1 points per doubling · 2.5× for -9.2
How to read this
Points per doubling is index points gained each time the bill doubles. Spending scales geometrically here — each setting tends to cost a multiple of the last rather than a fixed amount more — so a per-doubling figure compares models whose price ranges are nothing alike. A high number means the dial is worth turning; a low one means you are paying for thinking that does not show up in the score.
Only models Artificial Analysis has run at more than one setting appear, because the cost figure comes from those runs. It is one publisher’s harness on one benchmark suite: a model that reasons usefully on work unlike theirs will look worse here than it is.
Cite as: BenchLeader, “What more thinking costs”, https://www.benchleader.com/effort, data as of 11 Oct 2026.