FireRedAudio: Unified 9B Audio Model for 1-Hour Understanding & Generation

pmttyji · reddit · 2026-08-22

FireRedTeam released FireRedAudio, a 9B-parameter general-purpose audio language model featuring decoupled continuous representations. It uses a single backbone for both understanding (Audio Encoder) and generation (RedAE pathway), supporting ASR, audio understanding, zero-shot TTS, Instruct TTS, and speech editing. It can handle recordings up to one hour long with precise temporal grounding.

Also released is FireRedTTS3, a unified speech system with a Base variant for zero-shot cloning across 24 languages and 21 Chinese dialects, and an Instruct variant for natural language voice design and editing. The models report leading WER/CER and speaker similarity scores.

Original post →

More from Multimodal

Multimodal channel →