Lech Mazur Creative Writing
Short-story writing judged by a panel of models.
As of 19 Sept 2026, GPT-5 leads Lech Mazur Creative Writing on BenchLeader with 86.0%, ahead of Kimi K2 at 85.6%, across 42 model configurations with a published result.
- Published by
- Lech Mazurdata via Epoch AI Benchmarking Hub
- Category
- Human preference
- Index weight
- Reference only
- Models
- 42
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
Models write short stories to constraints and a panel of judge models scores them on craft criteria.
How it is scored
Mean judge score, rescaled to percent.
What to keep in mind
Judged by models, so it reflects their taste; widely followed for creative writing.
- 1GPT-5 (medium)86.0%
- 2Kimi K285.6%
- 3Claude Opus 4.184.7%
- 4o3-pro (medium)84.4%
- 5o3 (medium)83.9%
- 6Gemini 2.5 Pro83.8%
- 7Claude Opus 483.6%
- 8GPT-5 mini (medium)83.1%
- 9Qwen3 235B A22B83.0%
- 10DeepSeek R183.0%
- 11Qwen3 235B A22B 2507 (thinking)82.4%
- 12DeepSeek R1 052881.9%
- 13GPT-4o81.8%
- 14Claude Sonnet 481.4%
- 15Claude 3.7 Sonnet81.1%
42 of 42
| # | ||||
|---|---|---|---|---|
| 1 | 86.0% | 58.5 | 2025-08-07 | |
| 2 | 85.6% | 51.0 | 2025-07-12 | |
| 3 | 84.7% | 51.9 | 2025-08-05 | |
| 4 | 84.4% | – | 2025-06-10 | |
| 5 | 83.9% | 53.7 | 2025-04-16 | |
| 6 | 83.8% | 54.5 | 2025-06-17 | |
| 7 | 83.6% | 52.9 | 2025-05-22 | |
| 8 | 83.1% | 54.6 | 2025-08-07 | |
| 9 | 83.0% | 49.6 | 2025-04-29 | |
| 10 | 83.0% | 48.7 | 2025-01-20 | |
| 11 | 82.4% | 51.1 | 2025-07-25 | |
| 12 | 81.9% | 50.2 | 2025-05-28 | |
| 13 | 81.8% | 44.3 | 2024-11-20 | |
| 14 | 81.4% | 49.7 | 2025-05-22 | |
| 15 | 81.1% | 50.2 | 2025-02-24 | |
| 16 | 80.3% | 46.6 | 2024-10-22 | |
| 17 | 80.2% | – | 2025-03-05 | |
| 18 | 79.9% | 39.5 | 2025-03-12 | |
| 19 | 77.3% | 45.5 | 2025-05-07 | |
| 20 | 77.1% | 46.9 | 2025-08-05 | |
| 21 | 76.9% | 58.5 | 2025-07-09 | |
| 22 | 76.5% | 48.4 | 2025-04-17 | |
| 23 | 76.4% | 51.0 | 2025-04-09 | |
| 24 | 75.6% | 50.2 | 2025-02-27 | |
| 25 | 75.3% | 48.9 | 2025-04-29 | |
| 26 | 75.0% | 52.4 | 2025-04-16 | |
| 27 | 73.8% | 45.2 | 2025-01-21 | |
| 28 | 73.5% | 50.3 | 2025-04-09 | |
| 29 | 73.5% | 40.3 | 2024-10-22 | |
| 30 | 73.4% | 52.1 | 2025-08-03 | |
| 31 | 72.9% | 50.9 | 2025-01-28 | |
| 32 | 71.5% | 43.7 | 2024-12-11 | |
| 33 | 70.2% | 48.4 | 2024-12-17 | |
| 34 | 69.0% | 41.0 | 2024-07-24 | |
| 35 | 67.2% | 36.2 | 2024-07-18 | |
| 36 | 64.9% | 44.6 | 2024-09-12 | |
| 37 | 63.6% | 42.2 | 2024-12-12 | |
| 38 | 62.6% | 39.6 | 2024-12-12 | |
| 39 | 62.0% | 43.4 | 2025-04-05 | |
| 40 | 61.7% | 47.7 | 2025-01-31 | |
| 41 | 61.5% | 45.6 | 2025-01-31 | |
| 42 | 60.5% | 41.4 | 2024-12-03 |
Cite as: BenchLeader, “Lech Mazur Creative Writing leaderboard”, https://www.benchleader.com/benchmarks/lech_mazur_writing, data as of 19 Sept 2026.
Lech Mazur Creative Writing: questions
- What does Lech Mazur Creative Writing measure?
- Models write short stories to constraints and a panel of judge models scores them on craft criteria. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads Lech Mazur Creative Writing?
- GPT-5 leads Lech Mazur Creative Writing with 86.0% as of 19 Sept 2026, ahead of Kimi K2 at 85.6%.
- How many models have Lech Mazur Creative Writing results?
- 42 model configurations have a Lech Mazur Creative Writing result on BenchLeader, all taken from Lech Mazur via Epoch AI Benchmarking Hub.
- Who runs Lech Mazur Creative Writing and how often is it updated?
- Lech Mazur Creative Writing is published by Lech Mazur. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does Lech Mazur Creative Writing count toward the BenchLeader Index?
- No. Lech Mazur Creative Writing is shown for reference but left out of the composite index.