Jev + WebMCP solves 100% of benchmark tasks at 112x lower cost than GPT-6 Astra

QuanquanGu · x · 2026-09-19

The WebMCP benchmark team reports that Jev paired with Mercury 2.5, a fast low-cost LLM, solved all 49 browser tasks at roughly 112x lower model cost than GPT-6 Astra with code execution, and 245x lower than Astra using screenshot-based computer use.

Jev's raw browser-control accuracy was unremarkable — only 25/49 tasks solved. Adding WebMCP nearly doubled that to 49/49 while cutting costs further. The team used Browser Use's open-source Ultrafast harness with reliability improvements.

Aditya Grover suspects Jev itself is a diffusion LLM: generating structured outputs like JSON in parallel resembles infilling while sampling from a dLLM, which may explain the strong showing.

Related event: Jev with WebMCP Solves All Browser Tasks at a Fraction of GPT-6 Astra Cost(3 posts)→

Original post →

More from coding & agent

coding & agent channel →