WindTunnel benchmark: WebMCP makes browser agents 2.5-7.5x faster, 3-47x cheaper

FinanceYF5 · x · 2026-09-20

WindTunnel, an open benchmark from nekuda-ai, compares how browser agents interact with websites. Key insight: the hard part is choosing the right next step, not generating tool arguments—WebMCP compresses clicks, forms and menus into one explicit tool call, shrinking the decision space. On the leaderboard Jev + Mercury 2.5 (WebMCP) tops at 96.5 (49/49 tasks), with GPT-5.6 Luna native WebMCP at 91.5. Versus the median screen-driving agent (44/49), WebMCP setups score 27-50% higher, finish 2.5-7.5x faster, and cost 3-47x less per task. Tests and methods are open-sourced and reproducible.

Related event: WindTunnel Benchmark: WebMCP Speeds Browser Agents 2.5-7.5x(2 posts)→

Original post →

More from coding & agent

coding & agent channel →