Xiaohongshu FireRedTTS3: Unified Speech Generation and Editing Model
jiqizhixin · x · 2026-08-26
Xiaohongshu's FireRedTeam presents FireRedTTS3, a unified model for cloning, designing, and editing speech across 24 languages and 21 Chinese dialects.
At its core is RedAE, a speech tokenizer that injects semantic meaning into acoustic representations from the start, eliminating the need for separate modules or multi-stage pipelines. The model handles zero-shot voice cloning from seconds of audio, text-described voice design, and precise speech editing where only the highlighted segment changes.
It achieved double first place in voice cloning accuracy and similarity (3.04% WER / 78.8%), double first in 24-language cloning (3.75% / 84.8%, top two in 22 of 24 languages), and leading scores on voice design and speech editing benchmarks.
More from Multimodal
- Use Spaces Media Extractor to Lock Video to Music Beat — techhalla · 2026-08-26
- Build Scenes and Place Characters Using Reference Images — techhalla · 2026-08-26
- Create Paper-Style Redneck Animation with NB2 and H3 — techhalla · 2026-08-26
- User tests AI-generated stadium broadcast shot, asks for feedback — Living-Specialist692 · 2026-08-26
- User shares AI-generated Gachapon video with a bleak comment — Independent-Arm-7397 · 2026-08-26
- AI Prompt Transforms Photos into Unique Motivational Posters — aziz4ai · 2026-08-26