Mercury 2.5 diffusion LLM launches at 80% off, p99 voice latency cut from minutes to 1s
StefanoErmon · x · 2026-09-09
Inception Labs' Mercury 2.5 diffusion LLM is now available via its API, OpenRouter, and Baseten, with 100M free tokens for new accounts. Launch pricing is 80% off at $0.04/M input and $0.15/M output tokens. OpenCall AI uses Mercury for voice agents on live patient calls: after switching from an AI inference chip provider, p50 latency fell below 200ms and p99 dropped from several minutes to one second, fitting multi-step reasoning inside a live call's latency budget.
Related event: Inception Launches Mercury 2.5, Its Strongest Diffusion LLM(4 posts)→
More from Models
- K2 Horizon open-sources six model scales; 0.9B posts 48.5 on AIME 2026 — kimmonismus · 2026-09-09
- IFM open-sources K2 Horizon: six models, 20T tokens each, and a public reward-hacking audit — kimmonismus · 2026-09-09
- IFM's K2 Horizon: six models from 0.9B to 375B with only 4B/23B active params — kimmonismus · 2026-09-09
- IFM's K2 Horizon aims to make the whole training pipeline inspectable — kimmonismus · 2026-09-09
- Dev launches aggregator site collecting all statements on OpenAI's claimed Navier–Stokes proof — NathanpmYoung · 2026-09-09
- kalomaze: the AI training 'flywheel' is diagnostics, not naive training on transcripts — kalomaze · 2026-09-09