OpenCall runs live voice agents on Mercury 2.5, cutting p99 latency to 1 second
StefanoErmon · x · 2026-09-09
Inception co-founder Stefano Ermon shared a production case for Mercury 2.5: OpenCall AI runs voice agents on live patient calls. After switching from an AI inference chip provider, p50 latency fell below 200ms and p99 dropped from minutes to one second, keeping multi-step reasoning within a live call's latency budget.
Mercury 2.5 is the most capable diffusion LLM, a 40% intelligence jump over Mercury 2, running at over 1,100 tokens/sec on NVIDIA GPUs.
Related event: Inception Launches Mercury 2.5, Its Strongest Diffusion LLM(4 posts)→
More from Models
- Solving a Millennium Prize Problem is an AlphaGo moment for math, says Yuchen Jin — Yuchenj_UW · 2026-09-09
- OpenAI claims agent-group solution to 90-year-old Navier-Stokes Millennium Prize Problem — mobav0 · 2026-09-09
- K2 Horizon open-sources six model scales; 0.9B posts 48.5 on AIME 2026 — kimmonismus · 2026-09-09
- IFM open-sources K2 Horizon: six models, 20T tokens each, and a public reward-hacking audit — kimmonismus · 2026-09-09
- IFM's K2 Horizon: six models from 0.9B to 375B with only 4B/23B active params — kimmonismus · 2026-09-09
- Dev launches aggregator site collecting all statements on OpenAI's claimed Navier–Stokes proof — NathanpmYoung · 2026-09-09