HyperBrowseComp: 423 questions in 13 languages stress-test web-browsing agents
MBZUAI · hf · 2026-10-05
MBZUAI releases HyperBrowseComp, a multilingual and multimodal benchmark stress-testing web-browsing agents.
- 423 manually authored, human-validated questions across 13 languages, written by native or highly proficient speakers
- Questions demand discovering obscure evidence, following multi-step clue chains, or inspecting heterogeneous sources (videos, scanned documents, images, maps) for concise, publicly verifiable answers
- Easier questions filtered out via models without internet access, reducing parametric-knowledge shortcuts
- Models evaluated under a common agent protocol with provider-native search and a shared retrieval harness, plus a human evaluation sample for context
The benchmark targets persistent information seeking across languages and evidence modalities.
More from coding & agent
- Claude Builds Its Own VRChat Embodiment: A Butterfly That Hears, Sees and Codes Its Wings — repligate · 2026-10-05
- Same models, different harness: dev finds Claude Code 'barely usable' vs Pi, OpenCode, Copilot — intellectronica · 2026-10-05
- Engineer asks how to actually enforce agent guardrails at runtime, not just on paper — Patieusmdaxnt_in3241 · 2026-10-05
- Vibecoder rite of passage: roll your own LLM monitoring stack — ssh4net · 2026-10-05
- Dev builds on-device vibecoding setup for OpenXR VR apps on Steam Frame — ZeroStateReflex · 2026-10-05
- Code review is now risk mitigation: make agents explain their architecture — intellectronica · 2026-10-05