HiL-Bench
Whether agents notice information gaps and ask clarifying questions. Scale AI.
- Published by
- Scale AI SEAL
- Category
- Agents & tools
- Index weight
- 0.5
- Models
- 17
- Data as of
- 9 Sept 2026
Scale AI SEAL Leaderboards.
- 1Claude Fable 5.161.5%
- 2Claude Opus 557.0%
- 3Claude Fable 556.3%
- 4GLM-5.243.7%
- 5Claude Opus 4.741.7%
- 6Gemini 3.8 Flash41.5%
- 7GPT-5.539.7%
- 8Claude Opus 4.638.3%
- 9Claude Opus 4.835.3%
- 10Gemini 3.1 Pro35.3%
- 11GPT-5.6 Sol32.3%
- 12Gemini 3.5 Flash27.7%
- 13Grok 4.2020.0%
- 14Kimi K2.618.7%
- 15GPT-5.49.7%
17 of 17
| # | |||||
|---|---|---|---|---|---|
| 1 | 61.5% | Claude Fable 5.1 | 70.0 | – | |
| 2 | 57.0% | Claude Opus 5 | 70.0 | – | |
| 3 | 56.3% | Claude Fable 5 | 70.4 | – | |
| 4 | 43.7% | GLM 5.2 | 57.4 | – | |
| 5 | 41.7% | Claude Opus 4.7 | 65.5 | – | |
| 6 | 41.5% | Gemini 3.8 Flash | 62.4 | – | |
| 7 | 39.7% | GPT-5.5 | 66.6 | – | |
| 8 | 38.3% | Claude Opus 4.6 | 63.1 | – | |
| 9 | 35.3% | Claude Opus 4.8 | 64.7 | – | |
| 10 | 35.3% | Gemini 3.1 Pro | 63.9 | – | |
| 11 | 32.3% | GPT 5.6 Sol | 65.3 | – | |
| 12 | 27.7% | Gemini 3.5 Flash | 60.7 | – | |
| 13 | 20.0% | Grok-4.20 | 57.9 | – | |
| 14 | 18.7% | Kimi-k2.6 | 60.6 | – | |
| 15 | 9.7% | GPT-5.4 | 63.4 | – | |
| 16 | 6.3% | Minimax-M2.5 | 53.5 | – | |
| 17 | 4.3% | – | 62.4 | – |