BenchLeader

τ²-Bench Banking (AA)

Artificial Analysis' τ²-Bench banking-domain run: a support agent using tools with a simulated customer.

As of 19 Sept 2026, Qwen3 8 leads τ²-Bench Banking (AA) on BenchLeader with 51.3%, ahead of Grok 4.6 at 50.7%, across 204 model configurations with a published result.

Published by
Artificial Analysis
Category
Agents & tools
Index weight
1.0
Models
204
Data as of
19 Sept 2026

Source: Artificial Analysis (artificialanalysis.ai). Data taken from the public leaderboard.

What the test looks like

The banking domain of τ²-bench: the model acts as a bank's support agent, follows policy, uses tools, and coordinates with a simulated customer who also acts.

How it is scored

Percent of conversations ending in the correct final state, run by Artificial Analysis.

What to keep in mind

One domain, with a simulated customer who is itself a model. Counts toward the agentic category since index v1.2, alongside the telecom run.

204 of 204
#
1Qwen3 8maxAlibaba51.3%66.5
2Grok 4.6highxAI50.7%64.5
3Muse Spark 1.3maxMeta50.5%69.3
4GLM-5.3maxZhipu AIopen ↗50.3%65.8
5Qwen3.8 2.4T A95BAlibabaopen ↗49.1%65.2
6Qwen3.8 27BxhighAlibabaopen ↗48.0%59.2
7Agnes 3.0 FlashSapiens AI47.6%62.3
8Qwen3.8 27BmediumAlibabaopen ↗47.4%54.6
9Claude Fable 5.1thinkingAnthropic47.2%71.3
10Muse Spark 1.3xhighMeta47.2%68.0
11GLM-5.3-FlashZhipu AIopen ↗47.2%63.9
12Kimi K3maxMoonshot AIopen ↗46.0%67.2
13Claude Fable 5.1xhighAnthropic45.8%71.6
14Gemini 3.8 FlashmediumGoogle45.8%64.9
15Qwen3.8-Flash-NextAlibabaopen ↗45.4%61.6
16Gemini 3.8 FlashhighGoogle45.0%64.5
17Claude Opus 5highAnthropic44.7%70.2
18GPT-5.6 SolmaxOpenAI44.3%68.8
19Grok 4.6mediumxAI44.3%66.2
20Claude Opus 5xhighAnthropic43.3%70.2
21Grok 4.6xhighxAI43.3%64.6
22Claude Fable 5.1highAnthropic43.1%72.0
23GPT-6 AstraxhighOpenAI43.1%71.1
24Claude Opus 5maxAnthropic42.1%69.9
25Grok 4.5highxAI42.1%61.0
26Kimi K3lowMoonshot AIopen ↗41.6%59.6
27GPT-6 AstramaxOpenAI41.4%71.8
28Claude Fable 5.1mediumAnthropic41.0%69.6
29Deepseek v4 Flash VisionmaxDeepSeekopen41.0%59.8
30GPT-5.6 TerramaxOpenAI40.2%65.0
31GPT-6 AstrahighOpenAI40.0%71.9
32GPT-5.4xhighOpenAI39.6%65.4
33DeepSeek V4 PromaxDeepSeekopen ↗39.6%64.2
34DeepSeek V4 FlashmaxDeepSeekopen ↗39.4%62.1
35GPT-5.5xhighOpenAI39.0%67.6
36Claude Fable 5.1lowAnthropic39.0%67.5
37Claude Opus 5mediumAnthropic38.6%67.3
38Ling 3.0 Flash FinUnknownopen38.6%55.4
39Claude Fable 5thinkingAnthropic38.1%70.9
40GPT-5.6 SolxhighOpenAI38.1%68.2
41Grok 4.6lowxAI38.1%61.0
42Claude Sonnet 5maxAnthropic37.3%60.5
43GPT-5.6 SolhighOpenAI36.7%68.0
44GPT-5.5highOpenAI36.7%67.0
45GPT-5.6 SolmediumOpenAI36.5%66.1
46Agnes 2.5 Pro BetaSapiens AI35.7%60.8
47GPT-6 AstramediumOpenAI35.5%69.8
48Gemini 3.7 FlashmediumGoogle35.5%64.7
49Motif 3Motif Technologiesopen35.3%59.2
50Muse Spark 1.2xhighMeta34.9%64.1
51Claude Opus 4.7maxAnthropic34.6%64.3
52GLM-5.2maxZhipu AIopen ↗34.6%63.9
53Claude Sonnet 4.6maxAnthropic34.4%57.6
54Ling 3.0 Flash VLAnt Groupopen ↗34.4%56.7
55Claude Opus 4.8maxAnthropic34.2%64.2
56K2 Horizon 375B A23BMBZUAIopen34.2%57.6
57Gemini 3.8 FlashlowGoogle33.2%60.7
58Gemini 3.7 FlashhighGoogle32.8%64.6
59Gemini 3.5 FlashhighGoogle32.2%63.6
60Qwen3.8 27BlowAlibabaopen ↗32.2%53.4

Cite as: BenchLeader, “τ²-Bench Banking (AA) leaderboard”, https://www.benchleader.com/benchmarks/aa_tau_banking, data as of 19 Sept 2026.

τ²-Bench Banking (AA): questions

What does τ²-Bench Banking (AA) measure?
The banking domain of τ²-bench: the model acts as a bank's support agent, follows policy, uses tools, and coordinates with a simulated customer who also acts. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads τ²-Bench Banking (AA)?
Qwen3 8 leads τ²-Bench Banking (AA) with 51.3% as of 19 Sept 2026, ahead of Grok 4.6 at 50.7%.
How many models have τ²-Bench Banking (AA) results?
204 model configurations have a τ²-Bench Banking (AA) result on BenchLeader, all taken from Artificial Analysis.
Who runs τ²-Bench Banking (AA) and how often is it updated?
τ²-Bench Banking (AA) is published by Artificial Analysis. BenchLeader re-reads the published results every morning and records the date each result was published.
Does τ²-Bench Banking (AA) count toward the BenchLeader Index?
Yes. τ²-Bench Banking (AA) contributes to the agents & tools category of the BenchLeader Index, normalised so that 50 is the average of the evaluated models.