WebRetriever Real-World Web Benchmark Released
机器之心 · wechat · 2026-07-16
Synced announced the opening of registrations for the WebRetriever Global Challenge, centering on an evaluation benchmark for WebAgents in real-world browser environments.
Event & Research Background
- Hosted by MingLue Technology, co-organized by Peking University, the Institute of Automation (CAS), the AI and Robotics Innovation Center of CAS Hong Kong Institute (AIRIS), and Synced
- Total prize pool of $15,000, open to individuals or teams regardless of nationality or institutional background
- The corresponding paper has been accepted by ECCV 2026
What WebRetriever Solves
The article points out that existing WebAgent benchmarks often rely on a small number of simulated or self-built sites, which differ vastly from the real, open internet. Furthermore, many evaluations only check if the "operation is correct" without systematically measuring whether the task was actually completed.
Key Data
- Covers 800 real online websites and 1,550 cross-industry tasks
- Spans 8 major domains including tech, finance, healthcare, education, and government
- The proprietary NavEval framework achieves a 91.2% consistency rate with human expert judgments, compared to roughly 81% for existing state-of-the-art methods
- Data shows: Even the best-performing single model has a basic navigation success rate of less than half, and an end-to-end full task completion rate of only about 20%
- Conclusion: "Arriving" does not equal "completing"
Registration & Resources
The article provides Octo registration details, invitation codes, quick terminal registration methods, and links to the paper, dataset, code, and leaderboard homepage.
More from coding & agent
- HeyGen adds a media-sourcing skill for coding agents with 75k images and 10k tracks — HeyGen · 2026-07-22
- Agent search bottlenecks are now about variance, not raw latency — rohanpaul_ai · 2026-07-22
- LangSmith adds tracing for Pipecat, LiveKit, OpenAI Realtime, and Gemini Live — LangChain · 2026-07-22
- An MCP server signs every AI agent tool call into a verifiable Merkle chain — Funky_Chicken_22 · 2026-07-22
- Annotated transcript of a Claude Code team interview is now available — trq212 · 2026-07-22
- Claude Code skill uses 10 Markdown rules to make outputs ADHD-friendly — alex_verem · 2026-07-22