HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents
To evaluate web-browsing agents on persistent information retrieval, the authors introduce HyperBrowseComp, a multilingual and multimodal browsing benchmark comprising 423 manually authored and human-validated questions across 13 languages. Questions require locating obscure evidence, following multi-step clue chains, or inspecting heterogeneous sources such as videos, scanned documents, images, or maps. Easier questions are filtered out using models without internet access to prevent reliance on parametric knowledge alone. The benchmark provides a challenging testbed evaluated through provider-native search and shared external retrieval harnesses under a common agent protocol.
HyperBrowseComp comprises 423 manually authored and human-validated questions across 13 languages.