BrowseComp-Plus
BrowseComp questions answered against a fixed 100k-document corpus instead of the live web, so the retriever and the model can be judged apart. Scores shown are on the BM25 baseline retriever, the one column nearly every submission runs. Model coverage stops in late 2025.
As of 21 Sept 2026, GPT-5 leads BrowseComp-Plus on BenchLeader with 57.6%, ahead of o3 at 50.5%, across 21 model configurations with a published result.
- Published by
- BrowseComp-Plusdata via BrowseComp-Plus
- Category
- Agents & tools
- Index weight
- Reference only
- Models
- 21
- Data as of
- 21 Sept 2026
BrowseComp-Plus (texttron/BrowseComp-Plus), results dataset Tevatron/BrowseComp-Plus-results.
- 1GPT-557.6%
- 2o350.5%
- 3gpt-oss-120b (high)43.6%
- 4Tongyi DeepResearch 30B-A3B36.9%
- 5GLM 4.733.3%
- 6GLM-4.627.7%
- 7gpt-oss-120b (medium)24.6%
- 8gpt-oss-20b (high)21.4%
- 9Gemini 2.5 Pro19.9%
- 10gpt-oss-20b (medium)16.9%
- 11Gemini 2.5 Flash16.3%
- 12Claude Opus 415.5%
- 13GPT-4.115.3%
- 14Claude Sonnet 414.7%
- 15Kimi K214.6%
21 of 21
| # | ||||
|---|---|---|---|---|
| 1 | 57.6% | 57.2 | 2025-08-08 | |
| 2 | 50.5% | 60.3 | 2025-08-08 | |
| 3 | 43.6% | 49.6 | 2025-12-31 | |
| 4 | 36.9% | – | 2025-12-27 | |
| 5 | 33.3% | 52.4 | 2025-12-26 | |
| 6 | 27.7% | 49.8 | 2025-10-13 | |
| 7 | 24.6% | – | 2025-08-08 | |
| 8 | 21.4% | 44.7 | 2025-08-08 | |
| 9 | 19.9% | 54.4 | 2025-08-08 | |
| 10 | 16.9% | – | 2025-08-08 | |
| 11 | 16.3% | 52.5 | 2025-08-08 | |
| 12 | 15.5% | 52.9 | 2025-08-08 | |
| 13 | 15.3% | 48.9 | 2025-08-08 | |
| 14 | 14.7% | 49.8 | 2025-08-08 | |
| 15 | 14.6% | 50.9 | 2025-08-16 | |
| 16 | Websailor 32BUnknownopen | 12.1% | – | 2025-09-28 |
| 17 | 9.8% | 45.0 | 2025-08-08 | |
| 18 | 8.6% | 50.2 | 2025-08-16 | |
| 19 | Searchr1 32BUnknownopen | 4.1% | – | 2025-08-08 |
| 20 | 4.0% | 44.5 | 2025-08-08 | |
| 21 | 3.6% | 52.1 | 2025-08-08 |
Cite as: BenchLeader, “BrowseComp-Plus leaderboard”, https://www.benchleader.com/benchmarks/browsecomp_plus, data as of 21 Sept 2026.
BrowseComp-Plus: questions
- What does BrowseComp-Plus measure?
- BrowseComp questions answered against a fixed 100k-document corpus instead of the live web, so the retriever and the model can be judged apart. Scores shown are on the BM25 baseline retriever, the one column nearly every submission runs. Model coverage stops in late 2025. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads BrowseComp-Plus?
- GPT-5 leads BrowseComp-Plus with 57.6% as of 21 Sept 2026, ahead of o3 at 50.5%.
- How many models have BrowseComp-Plus results?
- 21 model configurations have a BrowseComp-Plus result on BenchLeader, all taken from BrowseComp-Plus via BrowseComp-Plus.
- Who runs BrowseComp-Plus and how often is it updated?
- BrowseComp-Plus is published by BrowseComp-Plus. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does BrowseComp-Plus count toward the BenchLeader Index?
- No. BrowseComp-Plus is shown for reference but left out of the composite index.