Inception ships Mercury 2.5: most capable diffusion LLM at 1,107 tokens/sec, 260K context

StefanoErmon · x · 2026-09-09

Inception Labs has released Mercury 2.5, which CEO Stefano Ermon calls the most capable diffusion LLM on the market and the largest diffusion language model ever trained.

Usage has grown over an order of magnitude since Mercury 2 launched, with dozens of enterprises in production across search, voice, and coding. NVIDIA endorsed the release. Inception also shared OpenCall running live patient-call voice agents on Mercury, cutting p99 latency to one second.

Related event: Inception Launches Mercury 2.5, Its Strongest Diffusion LLM(4 posts)→

Original post →

More from Infra

Infra channel →