First WebMCP benchmark: 3-5x faster, up to 23x cheaper, +11.6% task success vs computer use
laparisa · x · 2026-09-06
nekuda released WindTunnel (webmcp.com), an open benchmark that for the first time compares WebMCP against computer-use, DOM+vision, and accessibility-tree browser agent interfaces across speed, cost, success rate, and token usage on 49 tasks.
Key findings:
- Every WebMCP config solved 48/49 tasks vs a median 43/49 (+11.6%) for screen-driving agents
- 3-5x faster (7.8s vs 28.1s median per task)
- 4-23x cheaper (0.6 cents vs 5.5 cents median per task)
- +39% higher median final score (91.2 vs 65.8)
Leaderboard highlights:
- GPT-5.6 Luna with native WebMCP tops at 98.4, finishing tasks in 5.7s at $0.002
- Sonnet 5 DOM+vision hit the highest attempt success (98.6%) but cost $0.21 and 64k tokens per task — 35x native WebMCP cost
- Claude Opus 5 computer use ranked last (56.6, 50.4s and $0.139 per task)
Final score weights success 60%, cost 20%, time 20%. The takeaway: structured interfaces letting agents invoke site capabilities dramatically outperform pixel/DOM-driven screen operation.
More from coding & agent
- Merge Agent Handler adds six connectors for calling third-party tools — shensi · 2026-09-06
- Community iterates on three.js jelly render: HDRI lighting and better jelly material — eschadiol · 2026-09-06
- "Total OpenAI victory" debate: coder says Cursor beats Codex harness by a wide margin — IndraVahan · 2026-09-06
- GLM Coding Plan ups Flash quotas: unlimited in ZCode, 2x elsewhere — pcuenq · 2026-09-06
- prime-agent 0.9.2 ships astra and fable 5.1 support, adds model price and usage view — xeophon · 2026-09-06
- Do production AI agents actually need an 'Agent SRE'? A developer asks to be proven wrong — Fantastic-Sleep-3352 · 2026-09-06