WebFovea takes 2nd in WebRetriever Challenge 2026: most agent failures live in the harness, not the model

Jiangang Han · hf · 2026-10-08

The team behind WebFovea, a vision-based web agent, scored 57.0/100 for 2nd place in the WebRetriever Challenge 2026 (Protocol III of the WebRetriever benchmark), where agents must operate live websites from an entry URL and return verifiable answers.

Key insight: a capable multimodal LLM is necessary but not sufficient — model decisions pass through the harness, and four things must go right at every step: parsing the reply into the intended action, the action taking effect, accurate reporting back, and showing the model what it needs. Many observed failures occurred at these stages rather than in model reasoning: a coordinate-space mismatch placed every click at 3/4 of intended coordinates, native dropdowns/iframes/text boxes failed silently, and self-generated chat-template tokens contaminated 4.9% of episodes.

Using the same model across all four submissions, the official hidden-set score rose from 31.0 to 57.0 purely through harness hardening. The writeup includes component evidence (with negative results), failure analysis, limitations, and a roadmap including per-step model routing. Code: https://github.com/jianganghan/WebFovea

Original post →

More from coding & agent

coding & agent channel →