Grok Voice goes live on fal: 0.70s latency, word-level timestamps, 2-minute voice cloning
SpaceXAI · x · 2026-09-17
xAI's Grok Voice speech model is now live on fal, answering in 0.70 seconds and often completing tool calls before a sentence ends.
Features include transcription with word-level timestamps, text-to-speech in 30+ voices across 25+ languages, and voice cloning from just two minutes of audio. The demo video's voices are all Grok Voice, targeting low-latency customer-support voice agents.
More from Multimodal
- Mistral OCR's Vik Paruchuri launches a new document-extraction benchmark — VikParuchuri · 2026-09-17
- The 2.5-hour fully AI-generated Odyssey movie is 2.5 hours too long — The Verge AI · 2026-09-17
- Maiden USA: an AI-generated sci-fi comedy short — Living_Operation4319 · 2026-09-17
- Training a GRU planner to guide flow diffusion improves motion continuity in a homebrew video model — pixlpa · 2026-09-17
- Cine Sueños trailer unveils a 10-film AI-generated narrative universe — DreamArchitect_ES · 2026-09-17
- Dev Shows Off AAA-Quality Game Content Made With Spawn in Under a Day — TAbrodi · 2026-09-17