BenchLeader

Lech Mazur Creative Writing

Short-story writing judged by a panel of models.

As of 19 Sept 2026, GPT-5 leads Lech Mazur Creative Writing on BenchLeader with 86.0%, ahead of Kimi K2 at 85.6%, across 42 model configurations with a published result.

Published by
Lech Mazurdata via Epoch AI Benchmarking Hub
Category
Human preference
Index weight
Reference only
Models
42
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

Models write short stories to constraints and a panel of judge models scores them on craft criteria.

How it is scored

Mean judge score, rescaled to percent.

What to keep in mind

Judged by models, so it reflects their taste; widely followed for creative writing.

42 of 42
#
1GPT-5mediumOpenAI86.0%58.52025-08-07
2Kimi K2Moonshot AIopen85.6%51.02025-07-12
3Claude Opus 4.1Anthropic84.7%51.92025-08-05
4o3-promediumOpenAI84.4%2025-06-10
5o3mediumOpenAI83.9%53.72025-04-16
6Gemini 2.5 ProGoogle83.8%54.52025-06-17
7Claude Opus 4Anthropic83.6%52.92025-05-22
8GPT-5 minimediumOpenAI83.1%54.62025-08-07
9Qwen3 235B A22BAlibabaopen83.0%49.62025-04-29
10DeepSeek R1DeepSeekopen83.0%48.72025-01-20
11Qwen3 235B A22B 2507thinkingAlibabaopen82.4%51.12025-07-25
12DeepSeek R1 0528DeepSeekopen81.9%50.22025-05-28
13GPT-4oOpenAI81.8%44.32024-11-20
14Claude Sonnet 4Anthropic81.4%49.72025-05-22
15Claude 3.7 SonnetAnthropic81.1%50.22025-02-24
16Claude 3.5 SonnetAnthropic80.3%46.62024-10-22
17QwQ-32BthinkingAlibabaopen80.2%2025-03-05
18Gemma 3 27BGoogleopen79.9%39.52025-03-12
19Mistral Medium 3Mistral AI77.3%45.52025-05-07
20gpt-oss-120bOpenAIopen77.1%46.92025-08-05
21Grok 4xAI76.9%58.52025-07-09
22Gemini 2.5 FlashthinkingGoogle76.5%48.42025-04-17
23Grok 3xAI76.4%51.02025-04-09
24GPT-4.5OpenAI75.6%50.22025-02-27
25Qwen3 30B A3BAlibabaopen75.3%48.92025-04-29
26o4-minimediumOpenAI75.0%52.42025-04-16
27Gemini 2.0 FlashthinkingGoogle73.8%45.22025-01-21
28Grok 3 minilowxAI73.5%50.32025-04-09
29Claude 3.5 HaikuAnthropic73.5%40.32024-10-22
30GLM-4.5Zhipu AIopen73.4%52.12025-08-03
31Qwen2.5-MaxmaxAlibaba72.9%50.92025-01-28
32Gemini 2.0 FlashGoogle71.5%43.72024-12-11
33o1mediumOpenAI70.2%48.42024-12-17
34Mistral Large 2Mistral AIopen69.0%41.02024-07-24
35GPT-4o miniOpenAI67.2%36.22024-07-18
36o1-minimediumOpenAI64.9%44.62024-09-12
37Grok 2xAIopen63.6%42.22024-12-12
38Phi-4Microsoftopen62.6%39.62024-12-12
39Llama 4 MaverickMetaopen62.0%43.42025-04-05
40o3-minihighOpenAI61.7%47.72025-01-31
41o3-minimediumOpenAI61.5%45.62025-01-31
42Nova ProAmazon60.5%41.42024-12-03

Cite as: BenchLeader, “Lech Mazur Creative Writing leaderboard”, https://www.benchleader.com/benchmarks/lech_mazur_writing, data as of 19 Sept 2026.

Lech Mazur Creative Writing: questions

What does Lech Mazur Creative Writing measure?
Models write short stories to constraints and a panel of judge models scores them on craft criteria. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads Lech Mazur Creative Writing?
GPT-5 leads Lech Mazur Creative Writing with 86.0% as of 19 Sept 2026, ahead of Kimi K2 at 85.6%.
How many models have Lech Mazur Creative Writing results?
42 model configurations have a Lech Mazur Creative Writing result on BenchLeader, all taken from Lech Mazur via Epoch AI Benchmarking Hub.
Who runs Lech Mazur Creative Writing and how often is it updated?
Lech Mazur Creative Writing is published by Lech Mazur. BenchLeader re-reads the published results every morning and records the date each result was published.
Does Lech Mazur Creative Writing count toward the BenchLeader Index?
No. Lech Mazur Creative Writing is shown for reference but left out of the composite index.