Xiaohongshu Open-Sources FireRedAudio, a 9B Unified Audio Language Model
aigclink · x · 2026-08-31
Xiaohongshu's FireRedTeam has open-sourced FireRedAudio, a general-purpose audio language model built on a shared 9B-parameter LLM backbone: an Audio Encoder handles understanding while a RedAE pathway handles generation, with the two representations decoupled yet sharing the same reasoning core.
A single model supports ASR, audio understanding, zero-shot TTS, instruct TTS, semantic/acoustic speech editing, and accurate temporal grounding over recordings up to one hour. In practice, a one-hour podcast can be transcribed into a timestamped professional show script, with voice cloning, text-designed new voices, and audio editing — output quality approaches manual interview-material editing.
Related event: Xiaohongshu Open-Sources FireRedAudio, a 9B Unified Audio-Language Model(2 posts)→
More from Multimodal
- Breeze-TTS-2 demo now available on Hugging Face — BreezeBlue · 2026-09-01
- 求助:Anima 模型训练正常但推理生成模糊 — a_throwawayorsmthn · 2026-09-01
- 求测:Ideogram 4 INT8 量化版在 3060 12G 上的表现 — WhyDoiHearBosssMusic · 2026-09-01
- Help: Swapping a Mustache Using a Reference Image in Flux Workflows — sadboi2021 · 2026-09-01
- Fal's new model heralds next chapter for Hollywood VFX — briannekimmel · 2026-09-01
- Sea Angels Generated in p5.js Using Claude Opus 5 — anselm · 2026-09-01