Mercury 2.5 Preview launches on OpenRouter with 1,107 tok/s speed
StefanoErmon · x · 2026-09-02
Inception AI released Mercury 2.5 Preview exclusively on OpenRouter. The model targets low latency, achieving 1,107 tokens per second via parallel token generation. It features tunable reasoning, parallel tool calls, and schema-aligned JSON output, designed for latency-sensitive workflows.
Related event: Inception AI's Mercury 2.5 Hits Over 1,100 Tokens Per Second(2 posts)→
More from Models
- Gemini agentic video understanding launches with a developer guide — osanseviero · 2026-09-02
- Gemini adds agentic video understanding, cutting token usage by 88% — osanseviero · 2026-09-02
- Gemini Adds Agentic Video Understanding, Cuts Token Usage by 88% — GoogleDeepMind · 2026-09-02
- Fable 5.1 spotted in Claude support docs, release appears imminent — kimmonismus · 2026-09-02
- Meta's Muse Voice Transcribe Balances Speed and Accuracy with Adaptive Delay — AIatMeta · 2026-09-02
- Fable and Sol Ultra models accused of subtle hallucinations — StewartalsopIII · 2026-09-02