Inception ships Mercury 2.5, a diffusion LLM claiming 40% higher intelligence at 1,107 tokens/sec
thione · x · 2026-09-15
- Inception (CEO Stefano Ermon) released Mercury 2.5, billed as the most capable and largest diffusion LLM ever trained, with 40% higher intelligence than Mercury 2, comparable to cost-optimized frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite and Claude Haiku 4.5.
- Key specs: 1,107 tokens/sec on commodity NVIDIA GPUs, 260K context, $0.20/M input and $0.75/M output (80% launch discount: $0.04/$0.15).
- Supports tunable reasoning, parallel tool calls, schema-aligned JSON; usage has grown over an order of magnitude since Mercury 2, and 2.5 was trained from production failure cases.
- The post also references DeepMind's AlphaGenome Atlas launch (covered separately).
More from Models
- Grok trains on user data by default; business plans can opt out — carlosdponx · 2026-09-15
- Bolt Forge launches free until Oct 14 with GLM, DeepSeek and Kimi plus up to 50x more usage — HeyAmit_ · 2026-09-15
- GPT-6 Astra tested on robot control: impressive on simple tasks, limited dexterity — DJiafei · 2026-09-15
- SOTA Inference Is Nearly Free for Consumers, So the Local-Model Trend May Reverse — mobileraj · 2026-09-15
- Cursor user switches to Claude Code, burns through quota by day 3 — jdluk87 · 2026-09-15
- Claude Max and Codex tiers are creating a computing power gap that locks out $20-budget newcomers — IndraVahan · 2026-09-15