WindTunnel benchmark: WebMCP makes browser agents 2.5-7.5x faster, 3-47x cheaper
FinanceYF5 · x · 2026-09-20
WindTunnel, an open benchmark from nekuda-ai, compares how browser agents interact with websites. Key insight: the hard part is choosing the right next step, not generating tool arguments—WebMCP compresses clicks, forms and menus into one explicit tool call, shrinking the decision space. On the leaderboard Jev + Mercury 2.5 (WebMCP) tops at 96.5 (49/49 tasks), with GPT-5.6 Luna native WebMCP at 91.5. Versus the median screen-driving agent (44/49), WebMCP setups score 27-50% higher, finish 2.5-7.5x faster, and cost 3-47x less per task. Tests and methods are open-sourced and reproducible.
Related event: WindTunnel Benchmark: WebMCP Speeds Browser Agents 2.5-7.5x(2 posts)→
More from coding & agent
- LLM attacks 800+ graph theory conjectures, yields 30+ full proofs — marc_lelarge · 2026-09-20
- GitHub tutorial with 15k stars teaches engineering reliable AI coding agents — tom_doerr · 2026-09-20
- Jev picks tools but can't generate text—Mercury 2.5 fills parameters at 1000+ tokens/sec — FinanceYF5 · 2026-09-20
- Slow prompt processing? Reddit user proposes pre-submitting agent overhead to warm the KV cache — butterfly_labs · 2026-09-20
- gemini-cli PR fixes silent remapping of explicit versioned model IDs — Pcmhacker-piro · 2026-09-20
- underclass pools multiple ChatGPT/Copilot subs behind one OpenAI-compatible local endpoint — airesearch12 · 2026-09-20