Benchmarking CPU-Based TTS Models via UTMOS
gvij · reddit · 2026-07-06
The author benchmarked four small TTS models on an Intel Xeon (4-core) CPU, scoring them with the objective UTMOS MOS metric: Kokoro 82M, Supertonic 3, Inflect-Nano (4.6M parameters), and Kyutai's newly released Pocket TTS (about 100M parameters). The results showed Kokoro 82M had the best audio quality (UTMOS approx. 4.44); thanks to its streaming LM architecture, Pocket TTS maintained a nearly flat RTF (0.69-0.76) regardless of text length variations, meaning inference cost scales linearly with output length and exhibits a uniquely flat scaling profile.
Related event: Kyutai Open-Sources Pocket-TTS for CPU-Based Voice Cloning(4 posts)→
More from Multimodal
- Storyboard-first workflows are making AI dance videos and influencers more consistent — aftahi_ai · 2026-07-22
- Interactive video should be judged by responsiveness, not just frame quality — Soggy_Limit8864 · 2026-07-22
- Runpod MCP and Claude help spin up image and video generation workflows — 802high · 2026-07-22
- Midjourney prompt turns a bee into a glitching pixel explosion — michaelrabone · 2026-07-22
- A physics reward can improve video generation without creating a real physics engine — Dapper-Drawer4546 · 2026-07-22
- HeyGen adds a media-sourcing skill for coding agents with 75k images and 10k tracks — HeyGen · 2026-07-22