OpenCall runs live voice agents on Mercury 2.5, cutting p99 latency to 1 second

StefanoErmon · x · 2026-09-09

Inception co-founder Stefano Ermon shared a production case for Mercury 2.5: OpenCall AI runs voice agents on live patient calls. After switching from an AI inference chip provider, p50 latency fell below 200ms and p99 dropped from minutes to one second, keeping multi-step reasoning within a live call's latency budget.

Mercury 2.5 is the most capable diffusion LLM, a 40% intelligence jump over Mercury 2, running at over 1,100 tokens/sec on NVIDIA GPUs.

Related event: Inception Launches Mercury 2.5, Its Strongest Diffusion LLM(4 posts)→

Original post →

More from Models

Models channel →