Mercury 2.5 Launches with 1,107 Tokens/sec Inference Speed
volokuleshov · x · 2026-09-02
Inception AI released Mercury 2.5, the next generation of its Mercury diffusion models, achieving 1,107 tokens/sec via parallel token generation. The update focuses on agentic workloads, featuring tunable reasoning, parallel tool calls, and schema-aligned JSON output. It is now live on OpenRouter.
Related event: Inception AI's Mercury 2.5 Hits Over 1,100 Tokens Per Second(2 posts)→
More from Models
- Gemini agentic video understanding launches with a developer guide — osanseviero · 2026-09-02
- Gemini adds agentic video understanding, cutting token usage by 88% — osanseviero · 2026-09-02
- Gemini Adds Agentic Video Understanding, Cuts Token Usage by 88% — GoogleDeepMind · 2026-09-02
- Fable 5.1 spotted in Claude support docs, release appears imminent — kimmonismus · 2026-09-02
- Meta's Muse Voice Transcribe Balances Speed and Accuracy with Adaptive Delay — AIatMeta · 2026-09-02
- Fable and Sol Ultra models accused of subtle hallucinations — StewartalsopIII · 2026-09-02