IFEval (HELM)
Strict instruction following on verifiable constraints. Run by Stanford CRFM's HELM.
As of 12 Sept 2026, Grok 3 mini leads IFEval (HELM) on BenchLeader with 95.1%, ahead of Grok 4 at 94.9%, across 68 model configurations with a published result.
- Published by
- HELM Capabilities
- Category
- Instruction following
- Index weight
- 1.0
- Models
- 68
- Data as of
- 12 Sept 2026
Stanford Center for Research on Foundation Models (crfm.stanford.edu), Apache-2.0 results.
What the test looks like
IFEval gives the model writing tasks with verifiable instructions attached: respond in exactly three paragraphs, avoid the letter e, end with a given phrase, and so on. Each instruction can be checked by a program, so no judge model is needed.
How it is scored
Strict accuracy: the share of instructions satisfied exactly, as run by Stanford CRFM's HELM in its fixed harness.
What to keep in mind
It checks whether rules were obeyed, not whether the writing was any good, and the constraints are artificial. Near saturation for frontier models, so it mainly separates the mid-field.
- 1Grok 3 mini95.1%
- 2Grok 494.9%
- 3GPT-5.193.5%
- 4GPT-5 nano93.2%
- 5o4-mini92.9%
- 6GPT-5 mini92.7%
- 7Claude Opus 491.8%
- 8Llama 4 Maverick90.8%
- 9GPT-4.1 mini90.4%
- 10Gemini 2.5 Flash89.8%
- 11Granite 4.0 Small89.0%
- 12Grok 388.4%
- 13Gemini 3 Pro87.6%
- 14Mistral Large 287.6%
- 15GPT-587.5%
| # | |||||
|---|---|---|---|---|---|
| 1 | 95.1% | – | 51.8 | – | |
| 2 | 94.9% | 0709 | 58.6 | – | |
| 3 | 93.5% | – | 59.4 | – | |
| 4 | 93.2% | – | 52.5 | – | |
| 5 | 92.9% | – | 57.2 | – | |
| 6 | 92.7% | – | 57.9 | – | |
| 7 | 91.8% | 20250514 | 52.2 | – | |
| 8 | 90.8% | Llama 4 Maverick (17Bx128E) Instruct FP8 | 44.2 | – | |
| 9 | 90.4% | – | 46.3 | – | |
| 10 | 89.8% | – | 50.0 | – | |
| 11 | 89.0% | IBM Granite 4.0 Small | – | – | |
| 12 | 88.4% | – | 51.0 | – | |
| 13 | 87.6% | – | 62.0 | – | |
| 14 | 87.6% | 2411 | 40.8 | – | |
| 15 | 87.5% | – | 60.8 | – | |
| 16 | 87.2% | – | – | – | |
| 17 | 86.9% | – | 61.6 | – | |
| 18 | 85.6% | 20241022 | 46.3 | – | |
| 19 | 85.0% | 20250929 | 53.7 | – | |
| 20 | 85.0% | Kimi K2 Instruct | 51.2 | – | |
| 21 | 84.9% | – | 56.1 | – | |
| 22 | 84.9% | IBM Granite 4.0 Micro | 35.9 | – | |
| 23 | 84.3% | – | 39.1 | – | |
| 24 | 84.1% | Gemini 2.0 Flash | 43.5 | – | |
| 25 | 84.0% | – | 55.1 | – | |
| 26 | 84.0% | – | 54.5 | – | |
| 27 | 83.9% | 20250514 | 50.5 | – | |
| 28 | 83.8% | – | 48.8 | – | |
| 29 | 83.7% | Gemini 1.5 Pro (002) | 44.8 | – | |
| 30 | 83.6% | – | 48.8 | – | |
| 31 | 83.5% | Qwen3 235B A22B Instruct 2507 FP8 | 54.1 | – | |
| 32 | 83.4% | 20250219 | 50.2 | – | |
| 33 | 83.2% | DeepSeek v3 | 45.2 | – | |
| 34 | 83.1% | Gemini 1.5 Flash (002) | 40.8 | – | |
| 35 | 82.4% | – | 47.7 | – | |
| 36 | 82.3% | Palmyra X5 | – | – | |
| 37 | 82.1% | Llama 3.1 Instruct Turbo (70B) | 41.1 | – | |
| 38 | 81.8% | Llama 4 Scout (17Bx16E) Instruct | 40.4 | – | |
| 39 | 81.7% | – | 43.8 | – | |
| 40 | 81.6% | Qwen3 235B A22B FP8 Throughput | 44.8 | – | |
| 41 | 81.5% | – | 41.1 | – | |
| 42 | 81.2% | – | 47.0 | – | |
| 43 | 81.1% | Llama 3.1 Instruct Turbo (405B) | 43.1 | – | |
| 44 | 81.0% | – | 51.2 | – | |
| 45 | 81.0% | – | 43.8 | – | |
| 46 | 80.6% | Qwen2.5 Instruct Turbo (72B) | 42.7 | – | |
| 47 | 80.3% | – | 45.8 | – | |
| 48 | 80.1% | 20251001 | 48.6 | – | |
| 49 | 79.3% | – | – | – | |
| 50 | 79.2% | 20241022 | 40.0 | – | |
| 51 | 78.4% | – | 49.9 | – | |
| 52 | 78.2% | – | 36.6 | – | |
| 53 | 78.0% | OLMo 2 32B Instruct March 2025 | – | – | |
| 54 | 77.6% | – | 39.4 | – | |
| 55 | 76.7% | – | – | – | |
| 56 | 76.0% | – | 39.8 | – | |
| 57 | 75.0% | 2503 | 40.3 | – | |
| 58 | 74.3% | Llama 3.1 Instruct Turbo (8B) | 36.3 | – | |
| 59 | 74.1% | Qwen2.5 Instruct Turbo (7B) | 40.1 | – | |
| 60 | 73.2% | – | 46.0 | – |
Cite as: BenchLeader, “IFEval (HELM) leaderboard”, https://www.benchleader.com/benchmarks/helm_ifeval, data as of 12 Sept 2026.
IFEval (HELM): questions
- What does IFEval (HELM) measure?
- IFEval gives the model writing tasks with verifiable instructions attached: respond in exactly three paragraphs, avoid the letter e, end with a given phrase, and so on. Each instruction can be checked by a program, so no judge model is needed. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads IFEval (HELM)?
- Grok 3 mini leads IFEval (HELM) with 95.1% as of 12 Sept 2026, ahead of Grok 4 at 94.9%.
- How many models have IFEval (HELM) results?
- 68 model configurations have a IFEval (HELM) result on BenchLeader, all taken from HELM Capabilities.
- Who runs IFEval (HELM) and how often is it updated?
- IFEval (HELM) is published by HELM Capabilities. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does IFEval (HELM) count toward the BenchLeader Index?
- Yes. IFEval (HELM) contributes to the instruction following category of the BenchLeader Index, normalised so that 50 is the average of the evaluated models.