Jev + WebMCP solves 100% of benchmark tasks at 112x lower cost than GPT-6 Astra
QuanquanGu · x · 2026-09-19
The WebMCP benchmark team reports that Jev paired with Mercury 2.5, a fast low-cost LLM, solved all 49 browser tasks at roughly 112x lower model cost than GPT-6 Astra with code execution, and 245x lower than Astra using screenshot-based computer use.
Jev's raw browser-control accuracy was unremarkable — only 25/49 tasks solved. Adding WebMCP nearly doubled that to 49/49 while cutting costs further. The team used Browser Use's open-source Ultrafast harness with reliability improvements.
Aditya Grover suspects Jev itself is a diffusion LLM: generating structured outputs like JSON in parallel resembles infilling while sampling from a dLLM, which may explain the strong showing.
Related event: Jev with WebMCP Solves All Browser Tasks at a Fraction of GPT-6 Astra Cost(3 posts)→
More from coding & agent
- willcb: in-context learning for personal continual learning needs neurosymbolic harnesses — willcb · 2026-09-19
- Online RL via task flywheels from live interaction traces is what actually works today — willcb · 2026-09-19
- Dev's take: Codex worth the $200 tier, Claude fine at $100, top-tier Cursor underused — lxfater · 2026-09-19
- AgentSky launches as an 'OpenRouter for agents': run 40+ coding agents in-browser and compare costs side by side — Aiden_Tech_Ai · 2026-09-19
- qwen-code SDK v0.1.13 fixes microcompaction to preserve prompt-cache reuse in long agent sessions — github-actions[bot] · 2026-09-19
- Agent engineering splits into three tiers: scripts, System One decision models, and reasoning models — arpit_bhayani · 2026-09-19