Gradium's TTS beta cuts time-to-first-audio from ~250ms to under 50ms

RemiCadene · x · 2026-09-23

Gradium AI released its TTS beta model with sub-50ms time-to-first-audio at comparable quality, while most TTS models sit above 100ms TTFA (original 250ms).

The team credits the gain to co-designing the model architecture with a custom inference stack, and is hiring engineers to run audio models on GPUs. The beta is available to try now.

Related event: Gradium TTS Beta Delivers Sub-50ms Time-to-First-Audio(2 posts)→

Original post →

More from Multimodal

Multimodal channel →