Mercury 2.5 LLM hits 770 tokens per second in speed benchmarks

Retro_Dev · hn · 2026-09-24

According to Artificial Analysis benchmarks, Mercury 2.5 delivers inference at 770 tokens per second, far above typical frontier-model output speeds. The Mercury line uses a diffusion-based text generation approach, and these numbers underscore the route's potential for low-latency inference.

Original post →

More from Infra

Infra channel →