Gradium claims fastest TTS yet with ~50ms time-to-first-audio, tops sub-100ms naturalness
mattturck · x · 2026-10-02
Voice AI startup Gradium announced its latest Text-to-Speech model with a time-to-first-audio of roughly 50ms — reportedly the lowest latency among frontier TTS models — and the highest naturalness score of any sub-100ms model on public Speko benchmarks. The company argues low latency doesn't require trading away robustness: the model keeps high accuracy on hard cases like phone numbers and alphanumeric text while improving naturalness and expressiveness. Gradium recently extended its seed round to $100 million with NVIDIA among new investors and opened a Bay Area office, offering TTS, speech-to-text, live translation, voice cloning and on-device TTS.
More from Infra
- Edge0 open-sources app layer: 35B model on-device with just 1–2.5GB memory — FinanceYF5 · 2026-10-02
- SpaceX Transporter-18 launches 130 payloads including Google's orbital AI chips — XFreeze · 2026-10-02
- Top 10% of firms capture 99.5% of model-serving spend, AI compute data shows — bendee983 · 2026-10-02
- Extropic founder teases 'Thermo RSI is coming' in cryptic thermodynamic computing hype post — beffjezos · 2026-10-02
- Microsoft Foundry puts GPT-6, Claude Opus 5.5 and Grok-4.6 all in one catalog with 1M contexts — mustafasuleyman · 2026-10-02
- Musk: "Orbital compute is gonna be a very big deal" as SpaceX eyes space-based energy for AI — DimaZeniuk · 2026-10-02