NetEase Releases Confucius4-TTS: 14-Language, Transcript-Free Zero-Shot Voice Cloning
_akhaliq · x · 2026-08-18
NetEase Youdao has published the Confucius4-TTS paper on arXiv and released an upgraded open-source model. It is a multilingual, cross-lingual zero-shot TTS system supporting 14 languages. A key innovation is its independence from audio prompt transcripts, addressing the limitation of untranscribed in-the-wild reference audio. The architecture features a two-stage design: an LLM-based text-to-semantic (T2S) module with a learnable speaker encoder, and a conditional flow-matching semantic-to-acoustic (S2A) module. It achieves a WER of 3.73% on the CV3-Eval cross-lingual benchmark and ranks first in human evaluations for speaker similarity on internal tests.
More from Multimodal
- Observation: Claude Opus 5 dominates the 3D demo scene — techartist_ · 2026-08-18
- Insect Reconstruction via Gaussian Splatting — janusch_patas · 2026-08-18
- Westlake University et al. propose Three-Body Scattering for single-step SOTA image generation — jiqizhixin · 2026-08-18
- Seeking AI tools for realistic product reviewer avatars — Mysterious_Level852 · 2026-08-18
- Workflow: Using Midjourney & GPT IMG 2 for consistent style — aziz4ai · 2026-08-18
- Karate Warrior: Combining Midjourney & GPT IMG for video — aziz4ai · 2026-08-18