Greek TTS from 3.5h Data: Deterministic Prompts Fix Speaker Drift, Near-Human MOS
kastnerkyle · x · 2026-09-13
An Interspeech 2026 paper shows how to build high-quality single-speaker TTS for a low-resource language.
- Data curation: WhisperX alignment and filtering turns audiobook recordings into TTS-ready data
- Base model: fine-tuned Parler-TTS (880M), whose pre-training encodes phonetic priors transferable to Greek
- Key finding: LLM-generated style prompts cause speaker drift at inference; replacing them with deterministic prompts resolves it
- LoRA anchoring: a speaker-specific LoRA trained on just 3.5h of single-speaker data (5% of parameters) anchors identity
- Results: WER 10.7% (2.9 above the ASR floor), MOS-I 4.00 vs 4.36 for human speech, speaker consistency MOS-C 4.24 vs 4.30
Conclusion: robust single-speaker Greek TTS is achievable with limited curated data.
More from Multimodal
- 100+ design iterations and 70+ AI video passes: behind the scenes of Lenny's Summit opening — lennysan · 2026-09-14
- Livebound: open-source local workspace for publishing AI images to CivitAI — MoonbearAIArt · 2026-09-14
- MiniMax H3 Turns a Drop of Coffee Into an Entire City in Hailuo Design Hub — umesh_ai · 2026-09-14
- "AI Can Never Produce This Quality" — Higgsfield Claps Back With "What If It Actually Can?" — _AustinCalvert_ · 2026-09-14
- Director Ruairi Robinson Drops AI-Generated Music Video Are We Not Men? — FudgeAllOfYous · 2026-09-14
- Creator Chains Seedream 2.5 with Flux 3 via Luma to Make a 40-Second AI Short — LudovicCreator · 2026-09-14