Inception ships Mercury 2.5: diffusion LLM with 40% intelligence jump at 1,100 tokens/sec

timshi_ai · x · 2026-09-09

Inception Labs launched Mercury 2.5, claiming the most capable diffusion LLM on the market: a 40% intelligence jump over Mercury 2, running over 1,100 tokens/sec on widely available NVIDIA GPUs. It's live on the Inception API, OpenRouter, and Baseten.

Details: 260K context, reasoning, tool use, and structured output; intro pricing of $0.04/1M input and $0.15/1M output (80% off); claimed quality matching speed-optimized autoregressive models like GPT-5.2 and Claude Sonnet 4.5. Also announced: Mercury Voice for voice agents (TTFT under 170ms) and Mercury Router for prompt-based model routing.

Related event: Inception Launches Mercury 2.5, Its Strongest Diffusion LLM(4 posts)→

Original post →

More from Models

Models channel →