BenchLeader

ITBench SRE (AA)

Site-reliability incident tasks in a live IT environment, run by Artificial Analysis.

As of 19 Sept 2026, GPT-5.6 Sol leads ITBench SRE (AA) on BenchLeader with 56.2%, ahead of GPT-5.6 Terra at 51.0%, across 33 model configurations with a published result.

Published by
Artificial Analysis
Category
Agents & tools
Index weight
Reference only
Models
33
Data as of
19 Sept 2026

Source: Artificial Analysis (artificialanalysis.ai). Data taken from the public leaderboard.

What the test looks like

ITBench SRE drops an agent into a running IT environment with an incident to diagnose and fix, in the style of on-call engineering work.

How it is scored

Percent of incidents resolved, run by Artificial Analysis.

What to keep in mind

Few models have results yet and each task is long, so the score is volatile.

33 of 33
#
1GPT-5.6 SolmaxOpenAI56.2%68.8
2GPT-5.6 TerramaxOpenAI51.0%65.0
3Kimi K3maxMoonshot AIopen ↗47.7%67.2
4Claude Opus 4.7maxAnthropic46.7%64.3
5GPT-5.5xhighOpenAI45.8%67.6
6GLM-5.2maxZhipu AIopen ↗42.7%63.9
7Qwen3 7maxAlibaba42.5%61.0
8Gemini 3.5 FlashhighGoogle40.4%63.6
9GPT-5.6 LunamaxOpenAI40.3%60.2
10GLM-5.1thinkingZhipu AIopen ↗40.3%57.8
11Claude Sonnet 4.6maxAnthropic39.8%57.6
12DeepSeek V4 PromaxDeepSeekopen ↗38.3%64.2
13MiMo-V2.5-ProXiaomiopen ↗38.2%59.2
14Gemma 4 31BthinkingGoogleopen ↗37.3%52.8
15Qwen3.5 27BthinkingAlibaba35.5%55.5
16GPT-5.4 minixhighOpenAI35.2%55.7
17Qwen3.5 397B-A17BthinkingAlibabaopen ↗34.1%56.1
18Grok 4.3highxAI32.7%58.2
19DeepSeek V4 FlashmaxDeepSeekopen ↗31.5%62.1
20Kimi K2.6Moonshot AIopen ↗31.2%60.5
21Gemini 3.1 ProGoogle30.3%63.9
22Step 3.7 FlashStepFunopen ↗30.3%52.6
23Claude Haiku 4.5thinkingAnthropic27.3%47.5
24MiniMax-M2.7MiniMaxopen ↗26.5%56.1
25GPT-5.4 nanoxhighOpenAI24.4%54.6
26Gemma 4 26B A4BthinkingGoogle23.6%50.7
27Qwen3.5 35B-A3BthinkingAlibaba21.5%53.5
28GPT-5.4 minino reasoningOpenAI18.6%44.6
29Grok 4.1no reasoningxAI17.9%43.0
30GPT-5.4 nanono reasoningOpenAI12.9%42.8
31gpt-oss-120bhighOpenAIopen5.7%49.7
32Nemotron 3 SuperthinkingNVIDIAopen ↗1.1%51.5
33Llama 3.3 70BMetaopen0.6%41.4

Cite as: BenchLeader, “ITBench SRE (AA) leaderboard”, https://www.benchleader.com/benchmarks/aa_itbench_sre, data as of 19 Sept 2026.

ITBench SRE (AA): questions

What does ITBench SRE (AA) measure?
ITBench SRE drops an agent into a running IT environment with an incident to diagnose and fix, in the style of on-call engineering work. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads ITBench SRE (AA)?
GPT-5.6 Sol leads ITBench SRE (AA) with 56.2% as of 19 Sept 2026, ahead of GPT-5.6 Terra at 51.0%.
How many models have ITBench SRE (AA) results?
33 model configurations have a ITBench SRE (AA) result on BenchLeader, all taken from Artificial Analysis.
Who runs ITBench SRE (AA) and how often is it updated?
ITBench SRE (AA) is published by Artificial Analysis. BenchLeader re-reads the published results every morning and records the date each result was published.
Does ITBench SRE (AA) count toward the BenchLeader Index?
No. ITBench SRE (AA) is shown for reference but left out of the composite index.