Mistral's Voxtral TTS Hits 70ms Latency but Stays Closed Source
shashib · x · 2026-08-06
Mistral is betting that voice will become the primary interface for enterprises to direct AI agents. At the Ai4 2026 conference, the company heavily showcased its audio model, Voxtral TTS.
Key metrics and features of the model include:
- Ultra-low Latency: Achieves 70ms latency on a 10-second voice sample.
- High Real-Time Factor: Generated speech arrives 9.7x faster than the audio's actual length.
- Zero-shot Cloning: Requires only 3 seconds of reference audio for voice cloning.
Despite the impressive technical specs, Voxtral TTS is notably the only Mistral model that is not fully open-source, raising questions about the company's licensing strategy and definition of 'open'.
More from Models
- 48 Hours of Heavy Use: DeepSeek Flash Proves Remarkably Stable — yacineMTB · 2026-08-06
- Qwen Max Open-Weights Controversy Highlights Corporate AI Governance — The AI Daily Brief · 2026-08-06
- Boson AI Launches Higgs Realtime Speech-to-Speech Model — smolix · 2026-08-06
- Together AI Launches Kimi K3 API, Leads Key Inference Benchmarks — togethercompute · 2026-08-06
- Testing Liquid AI's LFM2.5 2.6B: 90t/s Speed but Fails at Tool Calling — curiousily_ · 2026-08-06
- DeepSeek's 2K-GPU model may beat Google's best, raising questions about GPU efficiency and Google's strategy — teortaxesTex · 2026-08-06