BenchLeader

BrowseComp-Plus

BrowseComp questions answered against a fixed 100k-document corpus instead of the live web, so the retriever and the model can be judged apart. Scores shown are on the BM25 baseline retriever, the one column nearly every submission runs. Model coverage stops in late 2025.

As of 21 Sept 2026, GPT-5 leads BrowseComp-Plus on BenchLeader with 57.6%, ahead of o3 at 50.5%, across 21 model configurations with a published result.

Published by
BrowseComp-Plusdata via BrowseComp-Plus
Category
Agents & tools
Index weight
Reference only
Models
21
Data as of
21 Sept 2026

BrowseComp-Plus (texttron/BrowseComp-Plus), results dataset Tevatron/BrowseComp-Plus-results.

21 of 21
#
1GPT-5OpenAI57.6%57.22025-08-08
2o3OpenAI50.5%60.32025-08-08
3gpt-oss-120bhighOpenAIopen ↗43.6%49.62025-12-31
4Tongyi DeepResearch 30B-A3BAlibabaopen36.9%2025-12-27
5GLM 4.7Zhipu AIopen ↗33.3%52.42025-12-26
6GLM-4.6Zhipu AIopen ↗27.7%49.82025-10-13
7gpt-oss-120bmediumOpenAIopen ↗24.6%2025-08-08
8gpt-oss-20bhighOpenAIopen ↗21.4%44.72025-08-08
9Gemini 2.5 ProGoogle19.9%54.42025-08-08
10gpt-oss-20bmediumOpenAIopen ↗16.9%2025-08-08
11Gemini 2.5 FlashGoogle16.3%52.52025-08-08
12Claude Opus 4Anthropic15.5%52.92025-08-08
13GPT-4.1OpenAI15.3%48.92025-08-08
14Claude Sonnet 4Anthropic14.7%49.82025-08-08
15Kimi K2Moonshot AIopen ↗14.6%50.92025-08-16
16Websailor 32BUnknownopen12.1%2025-09-28
17gpt-oss-120blowOpenAIopen ↗9.8%45.02025-08-08
18DeepSeek R1 0528DeepSeekopen ↗8.6%50.22025-08-16
19Searchr1 32BUnknownopen4.1%2025-08-08
20gpt-oss-20blowOpenAIopen ↗4.0%44.52025-08-08
21Qwen3 32BAlibabaopen ↗3.6%52.12025-08-08

Cite as: BenchLeader, “BrowseComp-Plus leaderboard”, https://www.benchleader.com/benchmarks/browsecomp_plus, data as of 21 Sept 2026.

BrowseComp-Plus: questions

What does BrowseComp-Plus measure?
BrowseComp questions answered against a fixed 100k-document corpus instead of the live web, so the retriever and the model can be judged apart. Scores shown are on the BM25 baseline retriever, the one column nearly every submission runs. Model coverage stops in late 2025. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads BrowseComp-Plus?
GPT-5 leads BrowseComp-Plus with 57.6% as of 21 Sept 2026, ahead of o3 at 50.5%.
How many models have BrowseComp-Plus results?
21 model configurations have a BrowseComp-Plus result on BenchLeader, all taken from BrowseComp-Plus via BrowseComp-Plus.
Who runs BrowseComp-Plus and how often is it updated?
BrowseComp-Plus is published by BrowseComp-Plus. BenchLeader re-reads the published results every morning and records the date each result was published.
Does BrowseComp-Plus count toward the BenchLeader Index?
No. BrowseComp-Plus is shown for reference but left out of the composite index.