BenchLeader

MirrorCode

Reproducing a program's behaviour from its outputs. Run by Epoch AI.

As of 19 Sept 2026, Claude Fable 5.1 leads MirrorCode on BenchLeader with 73.3%, ahead of Claude Fable 5 at 63.9%, across 8 model configurations with a published result.

Published by
Epoch AI Benchmarking Hub
Category
Coding
Index weight
Reference only
Models
8
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

The model sees a program's behaviour and must write code that mirrors it exactly.

How it is scored

Percent of programs matched, run by Epoch AI.

What to keep in mind

Very few models measured so far.

8 of 8
#
1Claude Fable 5.1highAnthropic73.3%72.02026-09-01
2Claude Fable 5highAnthropic63.9%64.02026-06-09
3GPT-6 AstrahighOpenAI46.7%71.92026-09-03
4Claude Opus 4.7highAnthropic31.1%62.42026-04-16
5GPT-5.6 SolhighOpenAI20.0%68.02026-07-09
6GPT-5.4highOpenAI15.6%59.02026-03-05
7GPT-5.5highOpenAI10.0%67.02026-04-23
8Gemini 3.1 ProhighGoogle8.9%60.72026-02-19

Cite as: BenchLeader, “MirrorCode leaderboard”, https://www.benchleader.com/benchmarks/mirrorcode, data as of 19 Sept 2026.

MirrorCode: questions

What does MirrorCode measure?
The model sees a program's behaviour and must write code that mirrors it exactly. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads MirrorCode?
Claude Fable 5.1 leads MirrorCode with 73.3% as of 19 Sept 2026, ahead of Claude Fable 5 at 63.9%.
How many models have MirrorCode results?
8 model configurations have a MirrorCode result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
Who runs MirrorCode and how often is it updated?
MirrorCode is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
Does MirrorCode count toward the BenchLeader Index?
No. MirrorCode is shown for reference but left out of the composite index.