CPU Benchmarks of 4 Open-Source TTS Models
Practical-Koala2831 · reddit · 2026-07-16
The author ran CPU benchmarks on Neo comparing 4 open-source TTS models: Kokoro 82M, Supertonic 3, Inflect-Nano, and Kyutai's Pocket TTS.
Test process: auto-writing code, designing methodology, building harness, running 180 timed generations, scoring with UTMOS, exporting charts and CSVs. Notable findings:
- Pocket TTS can clone voices on CPU with just 5 seconds of reference audio, no GPU or fine-tuning required
- Kokoro 82M sounds most natural subjectively, achieving 1.5x real-time
- Supertonic 3 quality mode 4x real-time
- UTMOS scored Inflect-Nano high, but subjectively the author found it "hoarse, robotic", showing metrics don't always align with perception
More from Multimodal
- Gemini Omni Flash turns a boat cabin into a cave in Flow by Google — chrisfirst · 2026-07-22
- A simple workflow to turn a photo into an image prompt using Gemini, Grok, or GPT Image — harshitagu72595 · 2026-07-22
- A Reddit user proposes a consistency LoRA to keep anime and game scenes visually stable — ThirdWorldBoy21 · 2026-07-22
- Hand-painted figurines run through Seedance look eerily alive — cocktailpeanut · 2026-07-22
- An AI agent-made bayou country music video is making the rounds on Reddit — LazyKaleidoscope4696 · 2026-07-22
- Testing Qwen 3 Image: Map Borders Shift Based on Prompts, Includes Chinese Labels — NirantK · 2026-07-22