BenchLeader

Vending-Bench 2

Running a simulated vending-machine business for a year.

As of 19 Sept 2026, Claude Opus 5 leads Vending-Bench 2 on BenchLeader with 11181.9, ahead of Claude Opus 4.7 at 10936.8, across 60 model configurations with a published result.

Published by
Andon Labsdata via Epoch AI Benchmarking Hub
Category
Agents & tools
Index weight
Reference only
Models
60
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

The agent manages a vending business over a simulated year: ordering, pricing, and dealing with suppliers; scored on final balance.

How it is scored

Final balance in simulated dollars, as published by Andon Labs.

What to keep in mind

A dollar figure, not a percent; long-horizon coherence is what it tests.

Top 15
  1. 1Claude Opus 511181.9
  2. 2Claude Opus 4.710936.8
  3. 3GPT-5.6 Sol9619.4
  4. 4Grok 4.69047.0
  5. 5GLM-5.28313.8
  6. 6GLM-5.38163.6
  7. 7Claude Opus 4.68017.6
  8. 8GPT-5.57523.8
  9. 9GPT-5.6 Terra7343.2
  10. 10Claude Sonnet 4.67204.1
  11. 11Muse Spark 1.16520.5
  12. 12Claude Sonnet 56377.7
  13. 13Kimi K2.66204.6
  14. 14GPT-5.46144.2
  15. 15GPT-5.3 Codex5940.1
60 of 60
#
1Claude Opus 5Anthropic11181.966.72026-07-24
2Claude Opus 4.7Anthropic10936.864.52026-04-16
3GPT-5.6 SolOpenAI9619.456.82026-07-09
4Grok 4.6xAI9047.059.32026-08-12
5GLM-5.2Zhipu AIopen ↗8313.852.12026-06-16
6GLM-5.3Zhipu AIopen ↗8163.62026-08-14
7Claude Opus 4.6Anthropic8017.663.72026-02-05
8GPT-5.5OpenAI7523.863.22026-04-23
9GPT-5.6 TerraOpenAI7343.22026-07-09
10Claude Sonnet 4.6Anthropic7204.159.02026-02-17
11Muse Spark 1.1Meta6520.565.22026-07-09
12Claude Sonnet 5Anthropic6377.757.22026-06-30
13Kimi K2.6Moonshot AIopen ↗6204.660.52026-04-20
14GPT-5.4OpenAI6144.259.22026-03-05
15GPT-5.3 CodexOpenAI5940.154.22026-02-05
16Claude Opus 4.8Anthropic5787.461.92026-05-28
17Claude Fable 5highAnthropic5680.364.02026-06-09
18GLM-5.1Zhipu AIopen ↗5634.456.92026-04-07
19Gemini 3 ProGoogle5478.261.12025-11-18
20Claude Fable 5.1Anthropic5421.664.52026-09-01
21Gemini 3.5 FlashGoogle5396.451.32026-05-19
22Kimi K3Moonshot AIopen ↗5165.064.22026-07-16
23Qwen3.6 PlusAlibabaopen5114.958.52026-03-31
24Kimi K2.7 CodeMoonshot AIopen ↗5082.955.92026-06-12
25Claude Fable 5lowAnthropic5018.52026-06-09
26Claude Opus 4.5Anthropic4967.158.22025-11-24
27Claude Fable 5maxAnthropic4966.665.82026-06-09
28Grok 4.20xAI4662.92026-02-17
29Claude Fable 5Anthropic4529.968.32026-06-09
30GLM-5Zhipu AIopen4432.154.72026-02-11
31Claude Fable 5mediumAnthropic4339.82026-06-09
32Qwen3 6maxAlibaba4254.263.32026-04-20
33GPT-5.6 LunaOpenAI4094.72026-07-09
34Grok 4.5xAI3887.460.32026-07-08
35Claude Sonnet 4.5Anthropic3838.754.32025-09-29
36Gemini 3.1 Pro CustomtoolsGoogle3774.32026-02-19
37Gemini 3 FlashGoogle3634.757.72025-12-17
38GPT-5.2OpenAI3591.358.42025-12-11
39DeepSeek V4 ProDeepSeekopen ↗3284.552.62026-04-24
40Claude Opus 4.8maxAnthropic2992.364.22026-05-28
41GLM-4.7Zhipu AIopen ↗2376.852.42025-12-22
42MiniMax-M3MiniMaxopen ↗2157.857.52026-06-01
43GPT-5.1OpenAI1473.457.22025-11-13
44Kimi K2.5Moonshot AIopen1198.553.92026-01-27
45Grok 4.1thinkingxAI1106.656.02025-11-19
46DeepSeek V3.2DeepSeekopen103456.02025-09-29
47Gemini 3.1 ProGoogle911.263.92026-02-19
48Gemini 2.5 ProGoogle573.654.52025-06-17
49Gemini 2.5 FlashGoogle548.852.52025-06-17
50Qwen3.5 FlashAlibaba462.752.52026-02-25
51Claude Haiku 4.5Anthropic458.948.12025-10-15
52Qwen3.5 27BAlibaba202.053.02026-02-24
53MiniMax-M2MiniMaxopen160.655.12025-10-27
54Qwen3 MaxmaxAlibaba71.653.52025-09-24
55Grok 4.3xAI35.350.52026-04-17
56Qwen3.5 PlusAlibaba0.52026-02-16
57Qwen3 235B A22B 2507thinkingAlibabaopen-11.351.12025-07-25
58gpt-oss-120bOpenAIopen-21.546.92025-08-05
59MiniMax-M2.5MiniMaxopen ↗-23.254.22026-02-12
60GPT-5 miniOpenAI-31.256.82025-08-07

Cite as: BenchLeader, “Vending-Bench 2 leaderboard”, https://www.benchleader.com/benchmarks/vending_bench_2, data as of 19 Sept 2026.

Vending-Bench 2: questions

What does Vending-Bench 2 measure?
The agent manages a vending business over a simulated year: ordering, pricing, and dealing with suppliers; scored on final balance. Scores are reported in score; higher is better.
Which AI model leads Vending-Bench 2?
Claude Opus 5 leads Vending-Bench 2 with 11181.9 as of 19 Sept 2026, ahead of Claude Opus 4.7 at 10936.8.
How many models have Vending-Bench 2 results?
60 model configurations have a Vending-Bench 2 result on BenchLeader, all taken from Andon Labs via Epoch AI Benchmarking Hub.
Who runs Vending-Bench 2 and how often is it updated?
Vending-Bench 2 is published by Andon Labs. BenchLeader re-reads the published results every morning and records the date each result was published.
Does Vending-Bench 2 count toward the BenchLeader Index?
No. Vending-Bench 2 is shown for reference but left out of the composite index.