BenchLeader

SREBench

Site-reliability incidents to diagnose and fix. Run by Vals AI.

As of 19 Sept 2026, GPT-6 Astra leads SREBench on BenchLeader with 56.9%, ahead of GPT-5.6 Sol at 30.5%, across 9 model configurations with a published result.

Published by
Vals AI
Category
Agents & tools
Index weight
Reference only
Models
9
Data as of
19 Sept 2026

Vals AI (vals.ai).

What the test looks like

An agent is given a broken system and must find and fix the fault, in the style of on-call engineering.

How it is scored

Percent of incidents resolved, run by Vals AI.

What to keep in mind

Few models measured; long tasks make scores volatile.

9 of 9
#
1GPT-6 AstramaxOpenAI56.9%71.82026-09-16
2GPT-5.6 SolmaxOpenAI30.5%68.82026-09-16
3Claude Fable 5.1Anthropic22.9%64.52026-09-16
4Claude Opus 5Anthropic12.2%66.72026-09-16
5Gemini 3.7 FlashhighGoogle4.6%64.62026-09-16
6GPT-5.5xhighOpenAI3.8%67.62026-09-16
7Grok 4.5highxAI0.8%61.02026-09-16
8DeepSeek V4.1 FlashhighDeepSeekopen ↗0.8%2026-09-16
9GLM-5.2maxZhipu AIopen ↗0.0%63.92026-09-16

Cite as: BenchLeader, “SREBench leaderboard”, https://www.benchleader.com/benchmarks/vals_srebench, data as of 19 Sept 2026.

SREBench: questions

What does SREBench measure?
An agent is given a broken system and must find and fix the fault, in the style of on-call engineering. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads SREBench?
GPT-6 Astra leads SREBench with 56.9% as of 19 Sept 2026, ahead of GPT-5.6 Sol at 30.5%.
How many models have SREBench results?
9 model configurations have a SREBench result on BenchLeader, all taken from Vals AI.
Who runs SREBench and how often is it updated?
SREBench is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
Does SREBench count toward the BenchLeader Index?
No. SREBench is shown for reference but left out of the composite index.