BenchLeader

BrowseComp

Hard questions whose answers are buried deep on the open web: the model has to browse, not recall. Published by OpenAI.

As of 21 Sept 2026, Gemini 3.1 Pro leads BrowseComp on BenchLeader with 31.2%, ahead of GPT-5 at 20.1%, across 39 model configurations with a published result.

Published by
OpenAIdata via Kaggle Benchmarks
Category
Agents & tools
Index weight
Reference only
Models
39
Data as of
21 Sept 2026

Kaggle (kaggle.com); each board credits its publishing lab.

39 of 39
#
1Gemini 3.1 ProGoogle31.2%63.82026-03-25
2GPT-5OpenAI20.1%57.22025-08-26
3Gemini 3 FlashGoogle16.9%57.62026-03-23
4Gemini 2.5 ProGoogle7.8%54.42025-06-13
5o1OpenAI6.3%53.82025-06-13
6GPT-4.1OpenAI3.5%48.92025-06-13
7o4-miniOpenAI3.3%57.42025-06-13
8DeepSeek R1 0528DeepSeekopen ↗3.1%50.22025-07-02
9Gemini 2.5 FlashGoogle2.5%52.52025-06-13
10Grok 3 miniSpaceXAI1.7%52.72025-06-13
11gpt-oss-120bOpenAIopen ↗1.6%46.92025-08-21
12o3-miniOpenAI1.5%50.42025-06-13
13Grok 3SpaceXAI1.3%51.02025-06-13
14DeepSeek V3DeepSeekopen ↗1.3%44.82025-06-13
15Gemini 3.1 Flash LiteGoogle1.1%54.52026-03-23
16Claude Opus 4Anthropic1.0%52.92025-06-13
17Claude Opus 4.1Anthropic1.0%51.92025-09-03
18Grok 2SpaceXAIopen1.0%42.22025-06-13
19o1-miniOpenAI0.9%46.82025-06-13
20Claude 3.5 SonnetAnthropic0.9%46.62025-06-13
21Claude 3.7 SonnetAnthropic0.9%50.22025-06-13
22GPT-4oOpenAI0.7%44.32025-06-13
23Claude Sonnet 4Anthropic0.6%49.82025-06-13
24GPT 3.5 Turbo 1106OpenAI0.6%39.22025-06-13
25GPT-4o miniOpenAI0.6%36.22025-06-13
26Gemini 2.0 FlashGoogle0.6%43.72025-06-13
27Mistral Large 2Mistral AIopen0.6%41.02025-07-02
28Gemma 3 27BGoogleopen ↗0.6%39.52025-06-13
29Qwen3 235B A22B 2507thinkingAlibabaopen ↗0.5%51.12025-08-21
30Gemma 3 12BGoogleopen ↗0.5%37.72025-06-13
31Mixtral 8x22BMistral AIopen0.5%2025-07-02
32Gemini 1.5 Pro 002Google0.4%44.92025-06-13
33Gemini 1.5 Flash 002Google0.4%40.82025-06-13
34Granite 4.0 H SmallIBMopen ↗0.4%40.62025-09-30
35Gemma 3 4BGoogleopen ↗0.4%35.12025-06-13
36Claude 3.5 HaikuAnthropic0.3%40.32025-06-13
37Ministral 8BMistral AIopen0.3%38.42025-07-02
38Ministral 3BMistral AIopen0.2%2025-07-02
39Gemini 1.5 Flash 8BGoogle0.1%40.12025-06-13

Cite as: BenchLeader, “BrowseComp leaderboard”, https://www.benchleader.com/benchmarks/kaggle_browsecomp, data as of 21 Sept 2026.

BrowseComp: questions

What does BrowseComp measure?
Hard questions whose answers are buried deep on the open web: the model has to browse, not recall. Published by OpenAI. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads BrowseComp?
Gemini 3.1 Pro leads BrowseComp with 31.2% as of 21 Sept 2026, ahead of GPT-5 at 20.1%.
How many models have BrowseComp results?
39 model configurations have a BrowseComp result on BenchLeader, all taken from OpenAI via Kaggle Benchmarks.
Who runs BrowseComp and how often is it updated?
BrowseComp is published by OpenAI. BenchLeader re-reads the published results every morning and records the date each result was published.
Does BrowseComp count toward the BenchLeader Index?
No. BrowseComp is shown for reference but left out of the composite index.