4 Open-Source CPU TTS Models Compared: Only Pocket TTS Supports Cloning
gvij · reddit · 2026-07-06
The author compared four open-source TTS models on a CPU (including Kyutai's newly released Pocket TTS). Pocket TTS is the first CPU open-source model to support zero-shot voice cloning (using 5 seconds of reference audio), featuring an MIT license and simple pip installation. The benchmark shows: Kokoro 82M has the highest quality (MOS 4.46) but lacks cloning support; Supertonic 3's 5-step method is relatively fast (4.2x), making it ideal for video dubbing; Inflect-Nano is the fastest but suffers from muffled audio; and Pocket TTS is the slowest (1.4x) but remains the only option capable of voice cloning.
Related event: Kyutai Open-Sources Pocket-TTS for CPU-Based Voice Cloning(4 posts)→
More from Multimodal
- H3 long-form video experiment: 7-hour render, int8, peaked at 192GB RAM — SIR_NVAX_A_LOT · 2026-09-11
- 3D ResNet Paper Crosses 3,000 Citations Eight Years After CVPR 2018 — HirokatuKataoka · 2026-09-11
- Non-coder builds full-featured Android ComfyUI client with ChatGPT, submits to Google Play — ComfierUI · 2026-09-11
- RunningHub open-sources H3Lightning, speeding up MiniMax H3 video generation 12x — 智东西 · 2026-09-11
- FastH3-Live hits 22fps: acceleration node benchmarks and the --vram-headroom trick — spartong945 · 2026-09-11
- Midjourney style code share: --sref 2912175708 — tisch_eins · 2026-09-11