Kyutai's 100M-parameter on-device PocketTTS is first speech model trained with drifting
kastnerkyle · x · 2026-10-02
Kyutai Labs released a technical blog on training PocketTTS, its 100M-parameter on-device TTS model, using "drifting," a recent one-step generative objective from Deng et al. They claim it's the first speech model — and first autoregressive model — trained this way, achieving under 1% WER with high quality and voice cloning.
More from Multimodal
- AI filmmaking debate shifts from best model to who owns the production pipeline — lmoroney · 2026-10-02
- Gradium claims fastest TTS yet with ~50ms time-to-first-audio, tops sub-100ms naturalness — mattturck · 2026-10-02
- Fizgig 6.8.1 ships Krea 2 slider LoRAs, Ultra mode, and Qwen Image 2.1 full fine-tuning — shootthesound · 2026-10-02
- Meta's MemLife: training-free egocentric video memory system gains 4.6-12.0%, plus RL-optimized writer MemOpt — meta · 2026-10-02
- Onepin launches as a post-TTS QA step that catches AI voice pronunciation errors — FellMentKE · 2026-10-02
- Seedance 2.0 Recreates a 1976 Church-Basement Social With Striking Vintage Detail — charis_ai · 2026-10-02