Mercury 2.5 Preview launches on OpenRouter with 1,107 tok/s speed

StefanoErmon · x · 2026-09-02

Inception AI released Mercury 2.5 Preview exclusively on OpenRouter. The model targets low latency, achieving 1,107 tokens per second via parallel token generation. It features tunable reasoning, parallel tool calls, and schema-aligned JSON output, designed for latency-sensitive workflows.

Related event: Inception AI's Mercury 2.5 Hits Over 1,100 Tokens Per Second(2 posts)→

Original post →

More from Models

Models channel →