Inception ships Mercury 2.5: diffusion LLM with 40% intelligence jump at 1,100 tokens/sec
timshi_ai · x · 2026-09-09
Inception Labs launched Mercury 2.5, claiming the most capable diffusion LLM on the market: a 40% intelligence jump over Mercury 2, running over 1,100 tokens/sec on widely available NVIDIA GPUs. It's live on the Inception API, OpenRouter, and Baseten.
Details: 260K context, reasoning, tool use, and structured output; intro pricing of $0.04/1M input and $0.15/1M output (80% off); claimed quality matching speed-optimized autoregressive models like GPT-5.2 and Claude Sonnet 4.5. Also announced: Mercury Voice for voice agents (TTFT under 170ms) and Mercury Router for prompt-based model routing.
Related event: Inception Launches Mercury 2.5, Its Strongest Diffusion LLM(4 posts)→
More from Models
- Commentary Claims OpenAI Frontier Labs Run Models a Year Ahead of the Public — rohanpaul_ai · 2026-09-09
- '50% of open problems just solved' — commentator marvels at frontier model progress — BorisMPower · 2026-09-09
- Cheap models via OpenRouter fall apart in agentic harnesses: GLM and DeepSeek can't match Claude — scottyLogJobs · 2026-09-09
- COLM 2026 paper: recent claims that LLMs can introspect don't meet the evidentiary bar — tallinzen · 2026-09-09
- Skeptic Dismisses OpenAI's Navier-Stokes Claim: 'It's Obviously in the Training Set' — adonis_singh · 2026-09-09
- Unverified screenshots show GPT-6 Astra debuting at No.2 on Agent Arena, 2 points behind Fable 5.1 — Gohab2001 · 2026-09-09