WildBench (HELM)
Real user tasks graded by a rubric-guided judge. Run by Stanford CRFM's HELM.
As of 12 Sept 2026, Qwen3 235B A22B 2507 leads WildBench (HELM) on BenchLeader with 86.6%, ahead of GPT-5.1 at 86.3%, across 68 model configurations with a published result.
- Published by
- HELM Capabilities
- Category
- Instruction following
- Index weight
- 0.5
- Models
- 68
- Data as of
- 12 Sept 2026
Stanford Center for Research on Foundation Models (crfm.stanford.edu), Apache-2.0 results.
What the test looks like
WildBench draws real tasks from users' conversations with chatbots, from coding help to creative writing, and grades responses with a judge model following task-specific checklists.
How it is scored
WB Score, a 0 to 100 rating from the checklist-guided judge, run by Stanford CRFM's HELM.
What to keep in mind
Judged by another model, so it inherits that judge's preferences; BenchLeader gives it half weight for that reason.
- 1Qwen3 235B A22B 250786.6%
- 2GPT-5.186.3%
- 3Kimi K286.2%
- 4o386.1%
- 5Gemini 3 Pro85.9%
- 6GPT-585.7%
- 7Gemini 2.5 Pro85.7%
- 8GPT-5 mini85.5%
- 9o4-mini85.4%
- 10Claude Sonnet 4.585.4%
- 11GPT-4.185.4%
- 12Claude Opus 485.2%
- 13Grok 384.9%
- 14gpt-oss-120b84.5%
- 15Claude Haiku 4.583.9%
| # | |||||
|---|---|---|---|---|---|
| 1 | 86.6% | Qwen3 235B A22B Instruct 2507 FP8 | 54.1 | – | |
| 2 | 86.3% | – | 59.4 | – | |
| 3 | 86.2% | Kimi K2 Instruct | 51.2 | – | |
| 4 | 86.1% | – | 61.6 | – | |
| 5 | 85.9% | – | 62.0 | – | |
| 6 | 85.7% | – | 60.8 | – | |
| 7 | 85.7% | – | 54.5 | – | |
| 8 | 85.5% | – | 57.9 | – | |
| 9 | 85.4% | – | 57.2 | – | |
| 10 | 85.4% | 20250929 | 53.7 | – | |
| 11 | 85.4% | – | 48.8 | – | |
| 12 | 85.2% | – | 56.1 | – | |
| 13 | 84.9% | – | 51.0 | – | |
| 14 | 84.5% | – | 48.8 | – | |
| 15 | 83.9% | 20251001 | 48.6 | – | |
| 16 | 83.8% | – | 55.1 | – | |
| 17 | 83.8% | – | 46.3 | – | |
| 18 | 83.3% | 20250514 | 52.2 | – | |
| 19 | 83.1% | DeepSeek v3 | 45.2 | – | |
| 20 | 82.8% | – | 49.9 | – | |
| 21 | 82.8% | Qwen3 235B A22B FP8 Throughput | 44.8 | – | |
| 22 | 82.8% | – | 43.8 | – | |
| 23 | 82.5% | 20250514 | 50.5 | – | |
| 24 | 81.8% | – | 43.8 | – | |
| 25 | 81.7% | – | 50.0 | – | |
| 26 | 81.4% | 20250219 | 50.2 | – | |
| 27 | 81.3% | Gemini 1.5 Pro (002) | 44.8 | – | |
| 28 | 81.1% | – | 39.1 | – | |
| 29 | 80.7% | – | 51.2 | – | |
| 30 | 80.6% | – | 52.5 | – | |
| 31 | 80.2% | Qwen2.5 Instruct Turbo (72B) | 42.7 | – | |
| 32 | 80.2% | – | – | – | |
| 33 | 80.1% | 2411 | 40.8 | – | |
| 34 | 80.0% | Llama 4 Maverick (17Bx128E) Instruct FP8 | 44.2 | – | |
| 35 | 80.0% | Gemini 2.0 Flash | 43.5 | – | |
| 36 | 79.7% | 0709 | 58.6 | – | |
| 37 | 79.2% | 20241022 | 46.3 | – | |
| 38 | 79.2% | Gemini 1.5 Flash (002) | 40.8 | – | |
| 39 | 79.1% | – | 36.6 | – | |
| 40 | 79.0% | – | 47.7 | – | |
| 41 | 78.9% | – | 47.0 | – | |
| 42 | 78.8% | – | 45.8 | – | |
| 43 | 78.8% | 2503 | 40.3 | – | |
| 44 | 78.3% | Llama 3.1 Instruct Turbo (405B) | 43.1 | – | |
| 45 | 78.3% | – | – | – | |
| 46 | 78.0% | Palmyra X5 | – | – | |
| 47 | 77.9% | Llama 4 Scout (17Bx16E) Instruct | 40.4 | – | |
| 48 | 77.7% | – | 41.1 | – | |
| 49 | 76.0% | 20241022 | 40.0 | – | |
| 50 | 75.8% | Llama 3.1 Instruct Turbo (70B) | 41.1 | – | |
| 51 | 75.0% | – | 39.4 | – | |
| 52 | 74.3% | – | 39.8 | – | |
| 53 | 74.1% | IBM Granite 3.3 8B Instruct | – | – | |
| 54 | 73.9% | IBM Granite 4.0 Small | – | – | |
| 55 | 73.7% | – | 46.0 | – | |
| 56 | 73.4% | OLMo 2 32B Instruct March 2025 | – | – | |
| 57 | 73.1% | Qwen2.5 Instruct Turbo (7B) | 40.1 | – | |
| 58 | 71.1% | Mixtral Instruct (8x22B) | 37.6 | – | |
| 59 | 68.9% | OLMo 2 13B Instruct November 2024 | – | – | |
| 60 | 68.6% | Llama 3.1 Instruct Turbo (8B) | 36.3 | – |
Cite as: BenchLeader, “WildBench (HELM) leaderboard”, https://www.benchleader.com/benchmarks/helm_wildbench, data as of 12 Sept 2026.
WildBench (HELM): questions
- What does WildBench (HELM) measure?
- WildBench draws real tasks from users' conversations with chatbots, from coding help to creative writing, and grades responses with a judge model following task-specific checklists. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads WildBench (HELM)?
- Qwen3 235B A22B 2507 leads WildBench (HELM) with 86.6% as of 12 Sept 2026, ahead of GPT-5.1 at 86.3%.
- How many models have WildBench (HELM) results?
- 68 model configurations have a WildBench (HELM) result on BenchLeader, all taken from HELM Capabilities.
- Who runs WildBench (HELM) and how often is it updated?
- WildBench (HELM) is published by HELM Capabilities. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does WildBench (HELM) count toward the BenchLeader Index?
- Yes. WildBench (HELM) contributes to the instruction following category of the BenchLeader Index, normalised so that 50 is the average of the evaluated models.