Mistral's Audio Model Talks Back in 70ms, Becomes Its First Non-Open Model
shashib · x · 2026-08-08
Mistral recently launched Voxtral TTS, its first text-to-speech model, aiming to make voice the primary interface for enterprises to direct AI agents.
Key data points from the analysis:
- 70ms: Claimed model latency for a 10-second voice sample.
- 9.7x: Claimed real-time factor, meaning generated speech arrives significantly faster than the audio's actual length.
- 3 sec: Reference audio required for zero-shot voice cloning.
Notably, Voxtral TTS is also the only model Mistral will not fully open-source, indicating a strategic shift in their licensing approach.
More from Companies & People
- Grok Ecosystem Accelerates: Top Image Model Released, Grok Build Hits 1M Visits — XFreeze · 2026-08-08
- Teacher-Free AI Private School Alpha School to Expand to 50 Campuses — 4KTV · 2026-08-08
- Report: US Government Holding Back OpenAI and Anthropic Models, Sparking Industry Pushback — haider1 · 2026-08-08
- Report: DeepMind Founder Hassabis Wanted to Leave; Google Moved Him to Chair to Prevent Stock Crash — firstadopter · 2026-08-08
- Maharashtra Govt Deploys Sarvam AI: 2,500 Officials to Get Indigenous AI Tools — RoboBalaji · 2026-08-08
- AI Brief: Google Leadership Shakeup, Meta's New Agents, Anthropic's Custom Chips — The AI Daily Brief · 2026-08-08