Xiaohongshu FireRedTTS3: Unified Speech Generation and Editing Model

jiqizhixin · x · 2026-08-26

Xiaohongshu's FireRedTeam presents FireRedTTS3, a unified model for cloning, designing, and editing speech across 24 languages and 21 Chinese dialects.

At its core is RedAE, a speech tokenizer that injects semantic meaning into acoustic representations from the start, eliminating the need for separate modules or multi-stage pipelines. The model handles zero-shot voice cloning from seconds of audio, text-described voice design, and precise speech editing where only the highlighted segment changes.

It achieved double first place in voice cloning accuracy and similarity (3.04% WER / 78.8%), double first in 24-language cloning (3.75% / 84.8%, top two in 22 of 24 languages), and leading scores on voice design and speech editing benchmarks.

Original post →

More from Multimodal

Multimodal channel →