Gradium claims fastest TTS yet with ~50ms time-to-first-audio, tops sub-100ms naturalness

mattturck · x · 2026-10-02

Voice AI startup Gradium announced its latest Text-to-Speech model with a time-to-first-audio of roughly 50ms — reportedly the lowest latency among frontier TTS models — and the highest naturalness score of any sub-100ms model on public Speko benchmarks. The company argues low latency doesn't require trading away robustness: the model keeps high accuracy on hard cases like phone numbers and alphanumeric text while improving naturalness and expressiveness. Gradium recently extended its seed round to $100 million with NVIDIA among new investors and opened a Bay Area office, offering TTS, speech-to-text, live translation, voice cloning and on-device TTS.

Original post →

More from Infra

Infra channel →