τ²-Bench Banking (AA)
Artificial Analysis' τ²-Bench banking-domain run: a support agent using tools with a simulated customer.
As of 19 Sept 2026, Qwen3 8 leads τ²-Bench Banking (AA) on BenchLeader with 51.3%, ahead of Grok 4.6 at 50.7%, across 204 model configurations with a published result.
- Published by
- Artificial Analysis
- Category
- Agents & tools
- Index weight
- 1.0
- Models
- 204
- Data as of
- 19 Sept 2026
Source: Artificial Analysis (artificialanalysis.ai). Data taken from the public leaderboard.
What the test looks like
The banking domain of τ²-bench: the model acts as a bank's support agent, follows policy, uses tools, and coordinates with a simulated customer who also acts.
How it is scored
Percent of conversations ending in the correct final state, run by Artificial Analysis.
What to keep in mind
One domain, with a simulated customer who is itself a model. Counts toward the agentic category since index v1.2, alongside the telecom run.
- 1Qwen3 8 (max)51.3%
- 2Grok 4.6 (high)50.7%
- 3Muse Spark 1.3 (max)50.5%
- 4GLM-5.3 (max)50.3%
- 5Qwen3.8 2.4T A95B49.1%
- 6Qwen3.8 27B (xhigh)48.0%
- 7Agnes 3.0 Flash47.6%
- 8Qwen3.8 27B (medium)47.4%
- 9Claude Fable 5.1 (thinking)47.2%
- 10Muse Spark 1.3 (xhigh)47.2%
- 11GLM-5.3-Flash47.2%
- 12Kimi K3 (max)46.0%
- 13Claude Fable 5.1 (xhigh)45.8%
- 14Gemini 3.8 Flash (medium)45.8%
- 15Qwen3.8-Flash-Next45.4%
| # | |||
|---|---|---|---|
| 1 | 51.3% | 66.5 | |
| 2 | 50.7% | 64.5 | |
| 3 | 50.5% | 69.3 | |
| 4 | 50.3% | 65.8 | |
| 5 | 49.1% | 65.2 | |
| 6 | 48.0% | 59.2 | |
| 7 | 47.6% | 62.3 | |
| 8 | 47.4% | 54.6 | |
| 9 | 47.2% | 71.3 | |
| 10 | 47.2% | 68.0 | |
| 11 | 47.2% | 63.9 | |
| 12 | 46.0% | 67.2 | |
| 13 | 45.8% | 71.6 | |
| 14 | 45.8% | 64.9 | |
| 15 | 45.4% | 61.6 | |
| 16 | 45.0% | 64.5 | |
| 17 | 44.7% | 70.2 | |
| 18 | 44.3% | 68.8 | |
| 19 | 44.3% | 66.2 | |
| 20 | 43.3% | 70.2 | |
| 21 | 43.3% | 64.6 | |
| 22 | 43.1% | 72.0 | |
| 23 | 43.1% | 71.1 | |
| 24 | 42.1% | 69.9 | |
| 25 | 42.1% | 61.0 | |
| 26 | 41.6% | 59.6 | |
| 27 | 41.4% | 71.8 | |
| 28 | 41.0% | 69.6 | |
| 29 | 41.0% | 59.8 | |
| 30 | 40.2% | 65.0 | |
| 31 | 40.0% | 71.9 | |
| 32 | 39.6% | 65.4 | |
| 33 | 39.6% | 64.2 | |
| 34 | 39.4% | 62.1 | |
| 35 | 39.0% | 67.6 | |
| 36 | 39.0% | 67.5 | |
| 37 | 38.6% | 67.3 | |
| 38 | Ling 3.0 Flash FinUnknownopen | 38.6% | 55.4 |
| 39 | 38.1% | 70.9 | |
| 40 | 38.1% | 68.2 | |
| 41 | 38.1% | 61.0 | |
| 42 | 37.3% | 60.5 | |
| 43 | 36.7% | 68.0 | |
| 44 | 36.7% | 67.0 | |
| 45 | 36.5% | 66.1 | |
| 46 | 35.7% | 60.8 | |
| 47 | 35.5% | 69.8 | |
| 48 | 35.5% | 64.7 | |
| 49 | 35.3% | 59.2 | |
| 50 | 34.9% | 64.1 | |
| 51 | 34.6% | 64.3 | |
| 52 | 34.6% | 63.9 | |
| 53 | 34.4% | 57.6 | |
| 54 | 34.4% | 56.7 | |
| 55 | 34.2% | 64.2 | |
| 56 | 34.2% | 57.6 | |
| 57 | 33.2% | 60.7 | |
| 58 | 32.8% | 64.6 | |
| 59 | 32.2% | 63.6 | |
| 60 | 32.2% | 53.4 |
Cite as: BenchLeader, “τ²-Bench Banking (AA) leaderboard”, https://www.benchleader.com/benchmarks/aa_tau_banking, data as of 19 Sept 2026.
τ²-Bench Banking (AA): questions
- What does τ²-Bench Banking (AA) measure?
- The banking domain of τ²-bench: the model acts as a bank's support agent, follows policy, uses tools, and coordinates with a simulated customer who also acts. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads τ²-Bench Banking (AA)?
- Qwen3 8 leads τ²-Bench Banking (AA) with 51.3% as of 19 Sept 2026, ahead of Grok 4.6 at 50.7%.
- How many models have τ²-Bench Banking (AA) results?
- 204 model configurations have a τ²-Bench Banking (AA) result on BenchLeader, all taken from Artificial Analysis.
- Who runs τ²-Bench Banking (AA) and how often is it updated?
- τ²-Bench Banking (AA) is published by Artificial Analysis. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does τ²-Bench Banking (AA) count toward the BenchLeader Index?
- Yes. τ²-Bench Banking (AA) contributes to the agents & tools category of the BenchLeader Index, normalised so that 50 is the average of the evaluated models.