Pocket TTS: Clone Any Voice in 5s on CPU
gvij · reddit · 2026-07-06
Kyutai's Pocket TTS is a streaming language model with roughly 100 million parameters. It uses the Mimi neural codec to generate audio tokens, enabling zero-shot voice cloning from just 5 seconds of reference audio without needing a GPU. It is MIT-licensed and easily installed via pip. The author compared it against three CPU TTS models: Kokoro, Supertonic, and Inflect-Nano. Across 180 timed tests, its latency remained nearly constant regardless of text length (RTF 0.69-0.76) and supported streaming output. Although it was the slowest of the four, it offered the most unique features.
Related event: Kyutai Open-Sources Pocket-TTS for CPU-Based Voice Cloning(4 posts)→
More from Multimodal
- Storyboard-first workflows are making AI dance videos and influencers more consistent — aftahi_ai · 2026-07-22
- Interactive video should be judged by responsiveness, not just frame quality — Soggy_Limit8864 · 2026-07-22
- Runpod MCP and Claude help spin up image and video generation workflows — 802high · 2026-07-22
- Midjourney prompt turns a bee into a glitching pixel explosion — michaelrabone · 2026-07-22
- A physics reward can improve video generation without creating a real physics engine — Dapper-Drawer4546 · 2026-07-22
- HeyGen adds a media-sourcing skill for coding agents with 75k images and 10k tracks — HeyGen · 2026-07-22