Mercury 2 Focuses on Low-Latency Inference

StefanoErmon · x · 2026-07-15

Mercury 2 has been described as "smart enough to reason, yet fast enough for voice call scenarios."

According to shared details, the model is the "world's first reasoning diffusion LLM," capable of completing full reasoning on standard NVIDIA GPUs in <300ms while achieving 1000+ tok/s.

The post also quoted the CEO of OpenCallAI, emphasizing that it delivers the necessary reasoning quality without sacrificing natural conversation latency.

Original post →

More from Models

Models channel →