Mercury 2.5 diffusion LLM launches at 80% off, p99 voice latency cut from minutes to 1s

StefanoErmon · x · 2026-09-09

Inception Labs' Mercury 2.5 diffusion LLM is now available via its API, OpenRouter, and Baseten, with 100M free tokens for new accounts. Launch pricing is 80% off at $0.04/M input and $0.15/M output tokens. OpenCall AI uses Mercury for voice agents on live patient calls: after switching from an AI inference chip provider, p50 latency fell below 200ms and p99 dropped from several minutes to one second, fitting multi-step reasoning inside a live call's latency budget.

Related event: Inception Launches Mercury 2.5, Its Strongest Diffusion LLM(4 posts)→

Original post →

More from Models

Models channel →