Mercury 2.5 Launches with 1,107 Tokens/sec Inference Speed

volokuleshov · x · 2026-09-02

Inception AI released Mercury 2.5, the next generation of its Mercury diffusion models, achieving 1,107 tokens/sec via parallel token generation. The update focuses on agentic workloads, featuring tunable reasoning, parallel tool calls, and schema-aligned JSON output. It is now live on OpenRouter.

Related event: Inception AI's Mercury 2.5 Hits Over 1,100 Tokens Per Second(2 posts)→

Original post →

More from Models

Models channel →