All resources / Technology & Ethics
BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
The PDF opens on the site where the paper is published.
- Authors
- Jason Wei, Zhiqing Sun, Spencer Papay, Scott McKinney, Jeffrey Han, Isa Fulford, Hyung Won Chung, Alex Tachard Passos, William Fedus, Amelia Glaese
- Venue
- arXiv (preprint — not peer-reviewed)
- Published
- April 16, 2025
- ID
- arXiv:2504.12516
Abstract
We present BrowseComp, a simple yet challenging benchmark for measuring the ability for agents to browse the web. BrowseComp comprises 1,266 questions that require persistently navigating the internet in search of hard-to-find, entangled information. Despite the difficulty of the questions, BrowseComp is simple and easy-to-use, as predicted answers are short and easily verifiable against reference answers. BrowseComp for browsing agents can be seen as analogous to how programming competitions are an incomplete but useful benchmark for coding agents. While BrowseComp sidesteps challenges of a true user query distribution, like generating long answers or resolving ambiguity, it measures the important core capability of exercising persistence and creativity in finding information.
The authors' own abstract, copied from arxiv.org.
How to cite it
Built from the details listed on this page. Check them against the paper before you submit.
Wei, J., Sun, Z., Papay, S., McKinney, S., Han, J., Fulford, I., Chung, H. W., Passos, A. T., Fedus, W., & Glaese, A. (2025). BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents. arXiv (preprint — not peer-reviewed). https://arxiv.org/abs/2504.12516
Wei, Jason, et al. "BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents." arXiv (preprint — not peer-reviewed), 2025, https://arxiv.org/abs/2504.12516.
@article{wei2025browsecomp,
title = {BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents},
author = {Jason Wei and Zhiqing Sun and Spencer Papay and Scott McKinney and Jeffrey Han and Isa Fulford and Hyung Won Chung and Alex Tachard Passos and William Fedus and Amelia Glaese},
journal = {arXiv (preprint — not peer-reviewed)},
year = {2025},
url = {https://arxiv.org/abs/2504.12516},
}What it is
OpenAI's benchmark for measuring how well agents browse the web: 1,266 questions that require persistently navigating the internet in search of hard-to-find, entangled information, with short answers that are easy to verify against a reference.
Topics
Added Oct 2, 2026 · 0 opens
