FireRedAudio Open-Sourced: One 9B Model That Listens, Reasons, Speaks and Edits Audio
aigclink · x · 2026-08-31
FireRedAudio is a general-purpose audio language model open-sourced by Xiaohongshu's FireRedTeam, built on a shared 9B-parameter LLM with decoupled continuous representations: an Audio Encoder pathway serves understanding while a RedAE-Patch pathway serves speech generation — claimed to be the first public design of its kind.
A single model covers ASR, audio understanding, zero-shot TTS, instruct TTS, semantic/acoustic speech editing, and accurate temporal grounding over recordings up to one hour. Official PyTorch code is available on GitHub (FireRedTeam/FireRedAudio).
Related event: Xiaohongshu Open-Sources FireRedAudio, a 9B Unified Audio-Language Model(2 posts)→
More from Multimodal
- Tip: Use Gemini for first pass, then Sol for detail correction — A_K_Nain · 2026-09-01
- Japanese Sports Challenge: 30s Photorealistic Video Gen — SimplyAnnisa · 2026-09-01
- Seedance 2.5 Upgrades AI Influencer Tools with Consistency and Storytelling — aftahi_ai · 2026-09-01
- RaVayan Character Generated with Krea_ai Seedance 2.5 — CurieuxExplorer · 2026-09-01
- AI sitcoms are now genuinely bingeable: "Bad Cat" hits episode four — venturetwins · 2026-09-01
- Spatial Studio adds Gaussian splats animation & 4K export — Scobleizer · 2026-09-01