Training-Free Accent Analogy Guidance Boosts Speaker Similarity in Cross-Lingual Voice Cloning
Yoomee Cho · hf · 2026-10-07
- In cross-lingual zero-shot TTS, reference-accent leakage contaminates cloned speech. Accent Analogy Guidance (AAG) is a training-free sampler term that isolates an accent direction from the model's own bilingual renders of one synthetic voice, cancelling voice identity and keeping only accent.
- The paper introduces ΔSIM, measuring speaker similarity above the identity-accent trade-off curve at equal accent. AAG beats the curve across four open TTS models: on OmniVoice, ΔSIM of +0.11 to +0.27 across three test sets; MaskGCT, CosyVoice 2, and F5-TTS also improve, while the method predicts where it won't help (X-Voice).
- Validated by a blind LLM accent judge, an LLM-free language-ID metric, and a 12-listener panel, all agreeing.
More from Multimodal
- LTX Director struggles: separating character actions from speech across multiple characters — salazar_slick · 2026-10-07
- VEDA Sparse Attention now available for MiniMax H3 in ComfyUI — robomar_ai_art · 2026-10-07
- VEDA Sparse Attention cuts MiniMax H3 video gen time in half in ComfyUI with no visible quality loss — robomar_ai_art · 2026-10-07
- a16z consumer AI overview turned into a 3-minute video with Manus — parker_lyman · 2026-10-07
- Blogger says Opus-generated video matches museum-grade immersive work that once cost $70K teams — oran_ge · 2026-10-07
- Super robot anime made with Kling 4 Flash looks straight out of a studio — Forsaken_Stuff_Ai · 2026-10-07