WebMCP benchmark: Jev + Mercury 2.5 solves 100% of tasks at 112x lower cost than GPT-6 Astra
hardimanjames · x · 2026-09-19
- A team ran Jev paired with Mercury 2.5 (a fast, low-cost LLM) on their WebMCP benchmark and it "basically broke" it: solving 100% of tasks at roughly 112x lower model cost than GPT-6 Astra with computer use and code execution, and about 245x lower vs Astra using screenshot-based computer use.
- Comparing Jev driving the browser directly (via Browser Use's open-source Ultrafast, with harness improvements) versus with WebMCP: browser control alone solved only 25/49 tasks; adding WebMCP nearly doubled that to 49/49 while cutting cost further.
Related event: Low-Cost Jev + Mercury Combo Solves WebMCP Benchmark at Fraction of Cost(2 posts)→
More from coding & agent
- Dev builds Concat, an open-source CapCut replacement, entirely with AI agents — JUB0T · 2026-09-19
- Dev builds Chrome extension with Jev to auto-tag sites and visualize browsing habits — pramodk73 · 2026-09-19
- One prompt, under $1: coding agent builds a full Django to-do app in 5 minutes — Al_Grigor · 2026-09-19
- One Generic Video Tool or Per-Provider Tools? Debating the Agent Tool Layer — ExcitingBison4616 · 2026-09-19
- Readback: free MIT VS Code extension reads Claude Code replies aloud via Speechify — shauntrennery · 2026-09-19
- GitHub Next open-sources LocalJev, a local Jev-compatible API built on oMLX and DiffusionGemma — gaganghotra_ · 2026-09-19