Kyutai's 100M-parameter on-device PocketTTS is first speech model trained with drifting

kastnerkyle · x · 2026-10-02

Kyutai Labs released a technical blog on training PocketTTS, its 100M-parameter on-device TTS model, using "drifting," a recent one-step generative objective from Deng et al. They claim it's the first speech model — and first autoregressive model — trained this way, achieving under 1% WER with high quality and voice cloning.

Original post →

More from Multimodal

Multimodal channel →