BenchLeader

WildBench (HELM)

Real user tasks graded by a rubric-guided judge. Run by Stanford CRFM's HELM.

As of 12 Sept 2026, Qwen3 235B A22B 2507 leads WildBench (HELM) on BenchLeader with 86.6%, ahead of GPT-5.1 at 86.3%, across 68 model configurations with a published result.

Published by
HELM Capabilities
Category
Instruction following
Index weight
0.5
Models
68
Data as of
12 Sept 2026

Stanford Center for Research on Foundation Models (crfm.stanford.edu), Apache-2.0 results.

What the test looks like

WildBench draws real tasks from users' conversations with chatbots, from coding help to creative writing, and grades responses with a judge model following task-specific checklists.

How it is scored

WB Score, a 0 to 100 rating from the checklist-guided judge, run by Stanford CRFM's HELM.

What to keep in mind

Judged by another model, so it inherits that judge's preferences; BenchLeader gives it half weight for that reason.

68 of 68
#
1Qwen3 235B A22B 2507Alibabaopen86.6%Qwen3 235B A22B Instruct 2507 FP854.1
2GPT-5.1OpenAI86.3%59.4
3Kimi K2Moonshot AIopen86.2%Kimi K2 Instruct51.2
4o3OpenAI86.1%61.6
5Gemini 3 ProGoogle85.9%62.0
6GPT-5OpenAI85.7%60.8
7Gemini 2.5 ProGoogle85.7%54.5
8GPT-5 miniOpenAI85.5%57.9
9o4-miniOpenAI85.4%57.2
10Claude Sonnet 4.5Anthropic85.4%2025092953.7
11GPT-4.1OpenAI85.4%48.8
12Claude Opus 4Anthropic85.2%56.1
13Grok 3xAI84.9%51.0
14gpt-oss-120bOpenAIopen84.5%48.8
15Claude Haiku 4.5Anthropic83.9%2025100148.6
16Claude Sonnet 4Anthropic83.8%55.1
17GPT-4.1 miniOpenAI83.8%46.3
18Claude Opus 4Anthropic83.3%2025051452.2
19DeepSeek V3DeepSeekopen83.1%DeepSeek v345.2
20DeepSeek R1 0528DeepSeekopen82.8%49.9
21Qwen3 235B A22BAlibabaopen82.8%Qwen3 235B A22B FP8 Throughput44.8
22GPT-4oOpenAI82.8%43.8
23Claude Sonnet 4Anthropic82.5%2025051450.5
24Gemini 2.5 Flash-LiteGoogle81.8%43.8
25Gemini 2.5 FlashGoogle81.7%50.0
26Claude 3.7 SonnetAnthropic81.4%2025021950.2
27Gemini 1.5 Pro 002Google81.3%Gemini 1.5 Pro (002)44.8
28GPT-4.1 nanoOpenAI81.1%39.1
29Qwen3-Next 80B A3BAlibabaopen80.7%51.2
30GPT-5 nanoOpenAI80.6%52.5
31Qwen2.5 72BAlibabaopen80.2%Qwen2.5 Instruct Turbo (72B)42.7
32Palmyra X 004Writer80.2%
33Mistral Large 2Mistral AIopen80.1%241140.8
34Llama 4 MaverickMetaopen80.0%Llama 4 Maverick (17Bx128E) Instruct FP844.2
35Gemini 2.0 FlashGoogle80.0%Gemini 2.0 Flash43.5
36Grok 4xAI79.7%070958.6
37Claude 3.5 SonnetAnthropic79.2%2024102246.3
38Gemini 1.5 Flash 002Google79.2%Gemini 1.5 Flash (002)40.8
39GPT-4o miniOpenAI79.1%36.6
40Gemini 2.0 Flash-LiteGoogle79.0%47.7
41GLM-4.5-AirZhipu AIopen78.9%47.0
42Nova PremierAmazon78.8%45.8
43Mistral Small 3.1Mistral AIopen78.8%250340.3
44Llama 3.1 405BMetaopen78.3%Llama 3.1 Instruct Turbo (405B)43.1
45Palmyra FinWriter78.3%
46Palmyra X5Writer78.0%Palmyra X5
47Llama 4 ScoutMetaopen77.9%Llama 4 Scout (17Bx16E) Instruct40.4
48Nova ProAmazon77.7%41.1
49Claude 3.5 HaikuAnthropic76.0%2024102240.0
50Llama 3.1 70BMetaopen75.8%Llama 3.1 Instruct Turbo (70B)41.1
51Amazon Nova LiteAmazon75.0%39.4
52Nova MicroAmazon74.3%39.8
53Granite 3.3 8BIBM74.1%IBM Granite 3.3 8B Instruct
54Granite 4.0 SmallIBM73.9%IBM Granite 4.0 Small
55gpt-oss-20bOpenAIopen73.7%46.0
56Olmo 2 32B March 2025Ai273.4%OLMo 2 32B Instruct March 2025
57Qwen2.5 7B InstructAlibabaopen73.1%Qwen2.5 Instruct Turbo (7B)40.1
58Mixtral 8x22B InstructMistral AIopen71.1%Mixtral Instruct (8x22B)37.6
59Olmo 2 13B November 2024Ai268.9%OLMo 2 13B Instruct November 2024
60Llama 3.1 8BMetaopen68.6%Llama 3.1 Instruct Turbo (8B)36.3

Cite as: BenchLeader, “WildBench (HELM) leaderboard”, https://www.benchleader.com/benchmarks/helm_wildbench, data as of 12 Sept 2026.

WildBench (HELM): questions

What does WildBench (HELM) measure?
WildBench draws real tasks from users' conversations with chatbots, from coding help to creative writing, and grades responses with a judge model following task-specific checklists. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads WildBench (HELM)?
Qwen3 235B A22B 2507 leads WildBench (HELM) with 86.6% as of 12 Sept 2026, ahead of GPT-5.1 at 86.3%.
How many models have WildBench (HELM) results?
68 model configurations have a WildBench (HELM) result on BenchLeader, all taken from HELM Capabilities.
Who runs WildBench (HELM) and how often is it updated?
WildBench (HELM) is published by HELM Capabilities. BenchLeader re-reads the published results every morning and records the date each result was published.
Does WildBench (HELM) count toward the BenchLeader Index?
Yes. WildBench (HELM) contributes to the instruction following category of the BenchLeader Index, normalised so that 50 is the average of the evaluated models.