Grok Voice goes live on fal: 0.70s latency, word-level timestamps, 2-minute voice cloning

SpaceXAI · x · 2026-09-17

xAI's Grok Voice speech model is now live on fal, answering in 0.70 seconds and often completing tool calls before a sentence ends.

Features include transcription with word-level timestamps, text-to-speech in 30+ voices across 25+ languages, and voice cloning from just two minutes of audio. The demo video's voices are all Grok Voice, targeting low-latency customer-support voice agents.

Original post →

More from Multimodal

Multimodal channel →