Xiaohongshu's FireRedTTS3 Unifies Speech Generation and Editing, Tops Benchmarks
机器之心 · wechat · 2026-08-21
The Xiaohongshu FireRed team released FireRedTTS3, a unified model for multilingual/dialectal cloning, natural language sound design, and precise speech editing.
Core Breakthrough: RedAE Representation
- Introduces a semantic teacher model to inject semantic information into continuous speech representations during training, mitigating error accumulation in autoregressive generation.
- Avoids acoustic detail loss from quantization, preserving fine timbral characteristics.
Architecture
- Built on an LLM-DiT framework with Aggregator, Backbone Transformer, and DiT modules.
- Two versions: FireRedTTS3-Base (cloning focus) and FireRedTTS3-Instruct (instruction-based design and editing).
Benchmark Results
- Seed-TTS-Eval: Ranks first in speaker similarity for zero-shot Chinese and English cloning.
- MiniMax-MLS-Test: Achieves 84.8% average similarity and 3.75% error rate across 24 languages.
- InstructTTSEval: Leads in sound design instruction following.
- Ming-Freeform-Audio-Edit: Outperforms in speech editing (acoustic and semantic), achieving precise edits without drift.
More from Multimodal
- Causal-rCM: Open-Source Recipe for Streaming Video Generation — chenhsuanlin · 2026-08-21
- Alibaba Releases HappyShrimp; Minimax Open-Sources Music3 — thursdai_pod · 2026-08-21
- Creator Uses Grok and Kling AI to Make Poetic Video — creatoroff · 2026-08-21
- Redditor Uses AI Video to Make a Literal Catfish: Cat Head, Fish Body — littleteckmonkey · 2026-08-21
- Best repeatable AI image gen workflow after testing 99% approaches — EXM7777 · 2026-08-21
- Grok generates assets and edits migration video via chat — JoeJustice · 2026-08-21