Inception ships Mercury 2.5: most capable diffusion LLM at 1,107 tokens/sec, 260K context
StefanoErmon · x · 2026-09-09
Inception Labs has released Mercury 2.5, which CEO Stefano Ermon calls the most capable diffusion LLM on the market and the largest diffusion language model ever trained.
- Quality: 40% intelligence jump over Mercury 2; comparable to cost-optimized frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5
- Speed: 1,107 tokens/sec on widely available NVIDIA GPUs
- Context: 260K tokens
- Price: $0.20/M input, $0.75/M output, 80% off at launch ($0.04/$0.15)
- Capabilities: tunable reasoning, parallel tool calls, schema-aligned JSON
Usage has grown over an order of magnitude since Mercury 2 launched, with dozens of enterprises in production across search, voice, and coding. NVIDIA endorsed the release. Inception also shared OpenCall running live patient-call voice agents on Mercury, cutting p99 latency to one second.
Related event: Inception Launches Mercury 2.5, Its Strongest Diffusion LLM(4 posts)→
More from Infra
- Uno speeds up Qwen3-8B 2.5x by using diffusion for parallel token drafting — rohanpaul_ai · 2026-09-09
- Oligopoly Equilibrium: why semiconductor markets settle at ~3 players — BenBajarin · 2026-09-09
- Deft Robotics launches unified deployment platform for physical AI — Scobleizer · 2026-09-09
- One user's local AI rig: DGX Spark bandwidth disappoints, RTX Pro 6000 looks like a steal — EAccelerate_42 · 2026-09-09
- Perplexity CEO: inference now served on NVLink Blackwells, Vera Rubin next — AravSrinivas · 2026-09-09
- TSMC to start High-NA EUV production in 2030 as ASML gains broader supply roadmap — BenBajarin · 2026-09-09