BenchLeader

IFEval (HELM)

Strict instruction following on verifiable constraints. Run by Stanford CRFM's HELM.

As of 12 Sept 2026, Grok 3 mini leads IFEval (HELM) on BenchLeader with 95.1%, ahead of Grok 4 at 94.9%, across 68 model configurations with a published result.

Published by
HELM Capabilities
Category
Instruction following
Index weight
1.0
Models
68
Data as of
12 Sept 2026

Stanford Center for Research on Foundation Models (crfm.stanford.edu), Apache-2.0 results.

What the test looks like

IFEval gives the model writing tasks with verifiable instructions attached: respond in exactly three paragraphs, avoid the letter e, end with a given phrase, and so on. Each instruction can be checked by a program, so no judge model is needed.

How it is scored

Strict accuracy: the share of instructions satisfied exactly, as run by Stanford CRFM's HELM in its fixed harness.

What to keep in mind

It checks whether rules were obeyed, not whether the writing was any good, and the constraints are artificial. Near saturation for frontier models, so it mainly separates the mid-field.

68 of 68
#
1Grok 3 minixAI95.1%51.8
2Grok 4xAI94.9%070958.6
3GPT-5.1OpenAI93.5%59.4
4GPT-5 nanoOpenAI93.2%52.5
5o4-miniOpenAI92.9%57.2
6GPT-5 miniOpenAI92.7%57.9
7Claude Opus 4Anthropic91.8%2025051452.2
8Llama 4 MaverickMetaopen90.8%Llama 4 Maverick (17Bx128E) Instruct FP844.2
9GPT-4.1 miniOpenAI90.4%46.3
10Gemini 2.5 FlashGoogle89.8%50.0
11Granite 4.0 SmallIBM89.0%IBM Granite 4.0 Small
12Grok 3xAI88.4%51.0
13Gemini 3 ProGoogle87.6%62.0
14Mistral Large 2Mistral AIopen87.6%241140.8
15GPT-5OpenAI87.5%60.8
16Palmyra X 004Writer87.2%
17o3OpenAI86.9%61.6
18Claude 3.5 SonnetAnthropic85.6%2024102246.3
19Claude Sonnet 4.5Anthropic85.0%2025092953.7
20Kimi K2Moonshot AIopen85.0%Kimi K2 Instruct51.2
21Claude Opus 4Anthropic84.9%56.1
22granite-4.0-microIBMopen84.9%IBM Granite 4.0 Micro35.9
23GPT-4.1 nanoOpenAI84.3%39.1
24Gemini 2.0 FlashGoogle84.1%Gemini 2.0 Flash43.5
25Claude Sonnet 4Anthropic84.0%55.1
26Gemini 2.5 ProGoogle84.0%54.5
27Claude Sonnet 4Anthropic83.9%2025051450.5
28GPT-4.1OpenAI83.8%48.8
29Gemini 1.5 Pro 002Google83.7%Gemini 1.5 Pro (002)44.8
30gpt-oss-120bOpenAIopen83.6%48.8
31Qwen3 235B A22B 2507Alibabaopen83.5%Qwen3 235B A22B Instruct 2507 FP854.1
32Claude 3.7 SonnetAnthropic83.4%2025021950.2
33DeepSeek V3DeepSeekopen83.2%DeepSeek v345.2
34Gemini 1.5 Flash 002Google83.1%Gemini 1.5 Flash (002)40.8
35Gemini 2.0 Flash-LiteGoogle82.4%47.7
36Palmyra X5Writer82.3%Palmyra X5
37Llama 3.1 70BMetaopen82.1%Llama 3.1 Instruct Turbo (70B)41.1
38Llama 4 ScoutMetaopen81.8%Llama 4 Scout (17Bx16E) Instruct40.4
39GPT-4oOpenAI81.7%43.8
40Qwen3 235B A22BAlibabaopen81.6%Qwen3 235B A22B FP8 Throughput44.8
41Nova ProAmazon81.5%41.1
42GLM-4.5-AirZhipu AIopen81.2%47.0
43Llama 3.1 405BMetaopen81.1%Llama 3.1 Instruct Turbo (405B)43.1
44Qwen3-Next 80B A3BAlibabaopen81.0%51.2
45Gemini 2.5 Flash-LiteGoogle81.0%43.8
46Qwen2.5 72BAlibabaopen80.6%Qwen2.5 Instruct Turbo (72B)42.7
47Nova PremierAmazon80.3%45.8
48Claude Haiku 4.5Anthropic80.1%2025100148.6
49Palmyra FinWriter79.3%
50Claude 3.5 HaikuAnthropic79.2%2024102240.0
51DeepSeek R1 0528DeepSeekopen78.4%49.9
52GPT-4o miniOpenAI78.2%36.6
53Olmo 2 32B March 2025Ai278.0%OLMo 2 32B Instruct March 2025
54Amazon Nova LiteAmazon77.6%39.4
55Palmyra MedWriter76.7%
56Nova MicroAmazon76.0%39.8
57Mistral Small 3.1Mistral AIopen75.0%250340.3
58Llama 3.1 8BMetaopen74.3%Llama 3.1 Instruct Turbo (8B)36.3
59Qwen2.5 7B InstructAlibabaopen74.1%Qwen2.5 Instruct Turbo (7B)40.1
60gpt-oss-20bOpenAIopen73.2%46.0

Cite as: BenchLeader, “IFEval (HELM) leaderboard”, https://www.benchleader.com/benchmarks/helm_ifeval, data as of 12 Sept 2026.

IFEval (HELM): questions

What does IFEval (HELM) measure?
IFEval gives the model writing tasks with verifiable instructions attached: respond in exactly three paragraphs, avoid the letter e, end with a given phrase, and so on. Each instruction can be checked by a program, so no judge model is needed. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads IFEval (HELM)?
Grok 3 mini leads IFEval (HELM) with 95.1% as of 12 Sept 2026, ahead of Grok 4 at 94.9%.
How many models have IFEval (HELM) results?
68 model configurations have a IFEval (HELM) result on BenchLeader, all taken from HELM Capabilities.
Who runs IFEval (HELM) and how often is it updated?
IFEval (HELM) is published by HELM Capabilities. BenchLeader re-reads the published results every morning and records the date each result was published.
Does IFEval (HELM) count toward the BenchLeader Index?
Yes. IFEval (HELM) contributes to the instruction following category of the BenchLeader Index, normalised so that 50 is the average of the evaluated models.