WebMCP Benchmark: Jev + Mercury 2.5 Solves 100% of Tasks at 245x Lower Cost Than Computer Use
hackgoofer · x · 2026-09-18
The WindTunnel WebMCP benchmark reports Jev + Mercury 2.5 solved all 49 browser tasks via WebMCP at 112x lower model cost than GPT-6 Astra with code execution, and 245x lower than screenshot-based computer use. WebMCP setups solve tasks 2.5–7.5x faster and 3–47x cheaper than screen-driving agents; adding WebMCP nearly doubled Jev's solved tasks from 25/49 to 49/49.
More from coding & agent
- Univer ships Office Harness: isolated worktrees let agents parallel-edit connected docs — Scobleizer · 2026-09-18
- ThursdAI: TypeSafe's Jev decision model hits 70-500ms at $42/1B input tokens, outputs free — thursdai_pod · 2026-09-18
- Computer-use agents: 273 tests across 23 assistants, Muse books a haircut by phone — thursdai_pod · 2026-09-18
- AI Worth Using Podcast and OpenClaw Launch Hackathon to Build Your Startup's First AI Hire — heyneighbor · 2026-09-18
- 54k-star proxy runs Claude Code, Codex and 8 more coding agents through 50 free providers — eyishazyer · 2026-09-18
- simslim: open-source tool runs more iOS simulators on one Mac by killing unneeded daemons — tom_doerr · 2026-09-18