Open-source NAR TTS scores WER 1.30 (EN) / 0.99 (ZH) on Seed-TTS, seeks cursed test cases
Important_Drag_6890 · reddit · 2026-09-24
The ViiTorVoice team reports benchmark results for their open-source NAR TTS model on Seed-TTS test sets: 1.30 English WER, 0.99 Chinese WER, and 0.734 English speaker similarity. Model and inference code are on GitHub. Noting that benchmarks cover only a clean slice of real usage, they are collecting harder cases — mixed Chinese-English text, unusual names, numbers/acronyms/CLI commands, tongue twisters, long-form narration, and weird punctuation — and inviting users to submit sentences that break TTS systems.
More from Models
- Rumor: OpenAI Plans a Staggering Number of Releases Today, New Agent Teased — iruletheworldmo · 2026-09-24
- MiniMax M3.1 (Space Bunny Alpha) Spotted Thinking in Token-Saving Caveman Mode — crusaderky · 2026-09-24
- GPT-6 Series Loses GPT-5.6's Relentless Task-Completion Trait — jdjohnson · 2026-09-24
- philschmid name-drops Gemini 3.8 Flash, says just use Gemini for multimodal understanding — _philschmid · 2026-09-24
- FLock's THIS/THAT 1.2 decision model beats Claude Opus 5 and GPT-5.6 with one forward pass — matlabulous · 2026-09-24
- User rant: Claude's over-filtering blocks legal fictional content, far stricter than ChatGPT — Dogbold · 2026-09-24