Xiaohongshu Open-Sources FireRedAudio, a 9B Unified Audio Language Model

aigclink · x · 2026-08-31

Xiaohongshu's FireRedTeam has open-sourced FireRedAudio, a general-purpose audio language model built on a shared 9B-parameter LLM backbone: an Audio Encoder handles understanding while a RedAE pathway handles generation, with the two representations decoupled yet sharing the same reasoning core.

A single model supports ASR, audio understanding, zero-shot TTS, instruct TTS, semantic/acoustic speech editing, and accurate temporal grounding over recordings up to one hour. In practice, a one-hour podcast can be transcribed into a timestamped professional show script, with voice cloning, text-designed new voices, and audio editing — output quality approaches manual interview-material editing.

Related event: Xiaohongshu Open-Sources FireRedAudio, a 9B Unified Audio-Language Model(2 posts)→

Original post →

More from Multimodal

Multimodal channel →