WebFovea takes 2nd in WebRetriever Challenge 2026: most agent failures live in the harness, not the model
Jiangang Han · hf · 2026-10-08
The team behind WebFovea, a vision-based web agent, scored 57.0/100 for 2nd place in the WebRetriever Challenge 2026 (Protocol III of the WebRetriever benchmark), where agents must operate live websites from an entry URL and return verifiable answers.
Key insight: a capable multimodal LLM is necessary but not sufficient — model decisions pass through the harness, and four things must go right at every step: parsing the reply into the intended action, the action taking effect, accurate reporting back, and showing the model what it needs. Many observed failures occurred at these stages rather than in model reasoning: a coordinate-space mismatch placed every click at 3/4 of intended coordinates, native dropdowns/iframes/text boxes failed silently, and self-generated chat-template tokens contaminated 4.9% of episodes.
Using the same model across all four submissions, the official hidden-set score rose from 31.0 to 57.0 purely through harness hardening. The writeup includes component evidence (with negative results), failure analysis, limitations, and a roadmap including per-step model routing. Code: https://github.com/jianganghan/WebFovea
More from coding & agent
- Splash 1.3.0 cuts local agent first-token time from 19s to 1s via SSD offloading — songhan_mit · 2026-10-08
- Antigravity 2 v2.21.1 & CLI updates ship full-text chat search and Automations tab — rseroter · 2026-10-08
- AI Builds an Email Inbox in an Afternoon, But Deliverability Is a Lifetime Ordeal — dosco · 2026-10-08
- Building simulated worlds is easy; running them on potatoes is not — gandamu_ml · 2026-10-08
- FrugaLLM: Free Open-Source App Auto-Routes Agents to the Smartest Free LLM Endpoints — KalKyl · 2026-10-08
- Microsoft updates WSL2 kernel to Linux 6.18.54.1 with DXGKRNL fence sharing — unixterminal · 2026-10-08