FireRedAudio: Unified 9B Audio Model for 1-Hour Understanding & Generation
pmttyji · reddit · 2026-08-22
FireRedTeam released FireRedAudio, a 9B-parameter general-purpose audio language model featuring decoupled continuous representations. It uses a single backbone for both understanding (Audio Encoder) and generation (RedAE pathway), supporting ASR, audio understanding, zero-shot TTS, Instruct TTS, and speech editing. It can handle recordings up to one hour long with precise temporal grounding.
Also released is FireRedTTS3, a unified speech system with a Base variant for zero-shot cloning across 24 languages and 21 Chinese dialects, and an Instruct variant for natural language voice design and editing. The models report leading WER/CER and speaker similarity scores.
More from Multimodal
- MiniMax-H3 video inpainting ported to diffusers modular blocks: 6-step subject swap — linoy_tsaban · 2026-08-22
- PixVerse gives away 6M credits for free AI video generation — aziz4ai · 2026-08-22
- 9-Minute Superhero Short Film Fully Generated with Seedance AI — Uerwol · 2026-08-22
- Community Speculation: Will FLUX 3 Go Open Source? — East_Shoe8814 · 2026-08-22
- Tested Gemini Omni Video-to-Video: SketchUp Preview Enhancement — aziz4ai · 2026-08-22
- Plamo 3D generation tool beta launches to bridge intent and execution — Scobleizer · 2026-08-22