4 Open-Source CPU TTS Models Compared: Only Pocket TTS Supports Cloning
gvij · reddit · 2026-07-06
The author compared four open-source TTS models on a CPU (including Kyutai's newly released Pocket TTS). Pocket TTS is the first CPU open-source model to support zero-shot voice cloning (using 5 seconds of reference audio), featuring an MIT license and simple pip installation. The benchmark shows: Kokoro 82M has the highest quality (MOS 4.46) but lacks cloning support; Supertonic 3's 5-step method is relatively fast (4.2x), making it ideal for video dubbing; Inflect-Nano is the fastest but suffers from muffled audio; and Pocket TTS is the slowest (1.4x) but remains the only option capable of voice cloning.
Related event: Kyutai Open-Sources Pocket-TTS for CPU-Based Voice Cloning(4 posts)→
More from Multimodal
- Storyboard-first workflows are making AI dance videos and influencers more consistent — aftahi_ai · 2026-07-22
- Interactive video should be judged by responsiveness, not just frame quality — Soggy_Limit8864 · 2026-07-22
- Runpod MCP and Claude help spin up image and video generation workflows — 802high · 2026-07-22
- Midjourney prompt turns a bee into a glitching pixel explosion — michaelrabone · 2026-07-22
- A physics reward can improve video generation without creating a real physics engine — Dapper-Drawer4546 · 2026-07-22
- HeyGen adds a media-sourcing skill for coding agents with 75k images and 10k tracks — HeyGen · 2026-07-22