HyperBrowseComp: a 423-question, 13-language stress test for web-browsing agents
zmkzmkz · x · 2026-10-06
Researchers released HyperBrowseComp (arXiv:2610.03574), a multilingual, multimodal stress-test benchmark for web-browsing agents.
- 423 handcrafted, human-validated questions across 13 languages, written by native speakers
- Answers are concise and publicly verifiable but require multi-step clue chains and obscure evidence in videos, scanned documents, images, or maps
- Easy questions filtered out via offline models to rule out parametric-knowledge answers
- Models evaluated under a common agent protocol, with human baseline on a sample
A co-author notes the German questions are among the hardest. Useful for benchmarking your favorite LLMs and agent harnesses.
More from coding & agent
- Tern, a Rust-native 'neoterminal', hits feature-complete with persistent agent sessions and 150ms startup — sull · 2026-10-06
- GTA 5 ported to WebAssembly with AI's help — 'the end of PC ports?' — pvncher · 2026-10-06
- How Rippling shipped production AI in 6 months with Deep Agents and LangSmith — LangChain · 2026-10-06
- COLM hosts Self-Improving Agents social with talks on Meta-Harness, EvoSkill and agent fragility — tuvllms · 2026-10-06
- Microsoft AI team shares talk on fine-tuning models for knowledge work like Excel — marlene_zw · 2026-10-06
- How do you change a production AI agent's authority without redeploying it? — BaraSlim · 2026-10-06