BrowseComp
Hard questions whose answers are buried deep on the open web: the model has to browse, not recall. Published by OpenAI.
As of 21 Sept 2026, Gemini 3.1 Pro leads BrowseComp on BenchLeader with 31.2%, ahead of GPT-5 at 20.1%, across 39 model configurations with a published result.
- Published by
- OpenAIdata via Kaggle Benchmarks
- Category
- Agents & tools
- Index weight
- Reference only
- Models
- 39
- Data as of
- 21 Sept 2026
Kaggle (kaggle.com); each board credits its publishing lab.
- 1Gemini 3.1 Pro31.2%
- 2GPT-520.1%
- 3Gemini 3 Flash16.9%
- 4Gemini 2.5 Pro7.8%
- 5o16.3%
- 6GPT-4.13.5%
- 7o4-mini3.3%
- 8DeepSeek R1 05283.1%
- 9Gemini 2.5 Flash2.5%
- 10Grok 3 mini1.7%
- 11gpt-oss-120b1.6%
- 12o3-mini1.5%
- 13Grok 31.3%
- 14DeepSeek V31.3%
- 15Gemini 3.1 Flash Lite1.1%
39 of 39
| # | ||||
|---|---|---|---|---|
| 1 | 31.2% | 63.8 | 2026-03-25 | |
| 2 | 20.1% | 57.2 | 2025-08-26 | |
| 3 | 16.9% | 57.6 | 2026-03-23 | |
| 4 | 7.8% | 54.4 | 2025-06-13 | |
| 5 | 6.3% | 53.8 | 2025-06-13 | |
| 6 | 3.5% | 48.9 | 2025-06-13 | |
| 7 | 3.3% | 57.4 | 2025-06-13 | |
| 8 | 3.1% | 50.2 | 2025-07-02 | |
| 9 | 2.5% | 52.5 | 2025-06-13 | |
| 10 | 1.7% | 52.7 | 2025-06-13 | |
| 11 | 1.6% | 46.9 | 2025-08-21 | |
| 12 | 1.5% | 50.4 | 2025-06-13 | |
| 13 | 1.3% | 51.0 | 2025-06-13 | |
| 14 | 1.3% | 44.8 | 2025-06-13 | |
| 15 | 1.1% | 54.5 | 2026-03-23 | |
| 16 | 1.0% | 52.9 | 2025-06-13 | |
| 17 | 1.0% | 51.9 | 2025-09-03 | |
| 18 | 1.0% | 42.2 | 2025-06-13 | |
| 19 | 0.9% | 46.8 | 2025-06-13 | |
| 20 | 0.9% | 46.6 | 2025-06-13 | |
| 21 | 0.9% | 50.2 | 2025-06-13 | |
| 22 | 0.7% | 44.3 | 2025-06-13 | |
| 23 | 0.6% | 49.8 | 2025-06-13 | |
| 24 | 0.6% | 39.2 | 2025-06-13 | |
| 25 | 0.6% | 36.2 | 2025-06-13 | |
| 26 | 0.6% | 43.7 | 2025-06-13 | |
| 27 | 0.6% | 41.0 | 2025-07-02 | |
| 28 | 0.6% | 39.5 | 2025-06-13 | |
| 29 | 0.5% | 51.1 | 2025-08-21 | |
| 30 | 0.5% | 37.7 | 2025-06-13 | |
| 31 | 0.5% | – | 2025-07-02 | |
| 32 | 0.4% | 44.9 | 2025-06-13 | |
| 33 | 0.4% | 40.8 | 2025-06-13 | |
| 34 | 0.4% | 40.6 | 2025-09-30 | |
| 35 | 0.4% | 35.1 | 2025-06-13 | |
| 36 | 0.3% | 40.3 | 2025-06-13 | |
| 37 | 0.3% | 38.4 | 2025-07-02 | |
| 38 | 0.2% | – | 2025-07-02 | |
| 39 | 0.1% | 40.1 | 2025-06-13 |
Cite as: BenchLeader, “BrowseComp leaderboard”, https://www.benchleader.com/benchmarks/kaggle_browsecomp, data as of 21 Sept 2026.
BrowseComp: questions
- What does BrowseComp measure?
- Hard questions whose answers are buried deep on the open web: the model has to browse, not recall. Published by OpenAI. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads BrowseComp?
- Gemini 3.1 Pro leads BrowseComp with 31.2% as of 21 Sept 2026, ahead of GPT-5 at 20.1%.
- How many models have BrowseComp results?
- 39 model configurations have a BrowseComp result on BenchLeader, all taken from OpenAI via Kaggle Benchmarks.
- Who runs BrowseComp and how often is it updated?
- BrowseComp is published by OpenAI. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does BrowseComp count toward the BenchLeader Index?
- No. BrowseComp is shown for reference but left out of the composite index.