Mercury 2 Focuses on Low-Latency Inference
StefanoErmon · x · 2026-07-15
Mercury 2 has been described as "smart enough to reason, yet fast enough for voice call scenarios."
According to shared details, the model is the "world's first reasoning diffusion LLM," capable of completing full reasoning on standard NVIDIA GPUs in <300ms while achieving 1000+ tok/s.
The post also quoted the CEO of OpenCallAI, emphasizing that it delivers the necessary reasoning quality without sacrificing natural conversation latency.
More from Models
- Moonshot’s Kimi K3 reaches #5 on MathArena as the top open model — xeophon · 2026-07-22
- Google launches Gemini 3.5 Flash Cyber for CodeMender, with limited access for governments — GoogleAI · 2026-07-22
- Kimi K3 tops Gemini 3.6 Flash on four shared public benchmarks — ChrisGPT · 2026-07-22
- Google’s year-long pause in new base-model pretraining draws sharp criticism — teortaxesTex · 2026-07-22
- Current setup is 8,192 input tokens and 2,048 output tokens, with 8k/512 next — TheZachMueller · 2026-07-22
- Kimi K3 feels slower than K2.7, but stronger on long coding jobs and refactoring — Far-Presence2711 · 2026-07-22