Benchmarking CPU-Based TTS Models via UTMOS
gvij · reddit · 2026-07-06
The author benchmarked four small TTS models on an Intel Xeon (4-core) CPU, scoring them with the objective UTMOS MOS metric: Kokoro 82M, Supertonic 3, Inflect-Nano (4.6M parameters), and Kyutai's newly released Pocket TTS (about 100M parameters). The results showed Kokoro 82M had the best audio quality (UTMOS approx. 4.44); thanks to its streaming LM architecture, Pocket TTS maintained a nearly flat RTF (0.69-0.76) regardless of text length variations, meaning inference cost scales linearly with output length and exhibits a uniquely flat scaling profile.
Related event: Kyutai Open-Sources Pocket-TTS for CPU-Based Voice Cloning(4 posts)→
More from Multimodal
- H3 long-form video experiment: 7-hour render, int8, peaked at 192GB RAM — SIR_NVAX_A_LOT · 2026-09-11
- 3D ResNet Paper Crosses 3,000 Citations Eight Years After CVPR 2018 — HirokatuKataoka · 2026-09-11
- Non-coder builds full-featured Android ComfyUI client with ChatGPT, submits to Google Play — ComfierUI · 2026-09-11
- RunningHub open-sources H3Lightning, speeding up MiniMax H3 video generation 12x — 智东西 · 2026-09-11
- FastH3-Live hits 22fps: acceleration node benchmarks and the --vram-headroom trick — spartong945 · 2026-09-11
- Midjourney style code share: --sref 2912175708 — tisch_eins · 2026-09-11