ANLI
Adversarially written natural-language inference. A 2019 benchmark, saturated for current models; kept for history.
As of 19 Sept 2026, GPT 3.5 Turbo 1106 leads ANLI on BenchLeader with 58.1%, ahead of Phi 3 Small at 58.1%, across 11 model configurations with a published result.
- Published by
- Epoch AI Benchmarking Hub
- Category
- Reasoning
- Index weight
- Reference only
- Models
- 11
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
Adversarially written natural-language inference.
How it is scored
Accuracy, as compiled by Epoch AI from published results.
What to keep in mind
Saturated: frontier models score near the ceiling, so it separates only older or smaller models, and results are largely self-reported by developers.
- 1GPT 3.5 Turbo 110658.1%
- 2Phi 3 Small58.1%
- 3Llama 3-8B57.3%
- 4phi-3-medium 14B (medium)55.8%
- 5Mixtral 8x7B55.2%
- 6Phi 3 Mini52.8%
- 7Gemma 7B48.7%
- 8Mistral 7B v0.347.1%
- 9Phi-242.5%
- 10Megatron-Turing NLG 530B39.7%
- 11InstructGPT 175B35.4%
| # | ||||
|---|---|---|---|---|
| 1 | 58.1% | 39.2 | 2023-11-06 | |
| 2 | 58.1% | – | 2024-04-23 | |
| 3 | 57.3% | 35.6 | 2024-04-18 | |
| 4 | 55.8% | – | 2024-04-23 | |
| 5 | 55.2% | 32.1 | 2023-12-11 | |
| 6 | 52.8% | 36.2 | 2024-04-23 | |
| 7 | 48.7% | – | 2024-02-21 | |
| 8 | 47.1% | 32.1 | 2023-09-27 | |
| 9 | 42.5% | – | 2023-12-12 | |
| 10 | 39.7% | – | 2021-10-11 | |
| 11 | 35.4% | – | 2022-01-27 |
Cite as: BenchLeader, “ANLI leaderboard”, https://www.benchleader.com/benchmarks/anli, data as of 19 Sept 2026.
ANLI: questions
- What does ANLI measure?
- Adversarially written natural-language inference. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads ANLI?
- GPT 3.5 Turbo 1106 leads ANLI with 58.1% as of 19 Sept 2026, ahead of Phi 3 Small at 58.1%.
- How many models have ANLI results?
- 11 model configurations have a ANLI result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
- Who runs ANLI and how often is it updated?
- ANLI is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does ANLI count toward the BenchLeader Index?
- No. ANLI is shown for reference but left out of the composite index.