MiDashengLM-Gen: Unified Audio Scene Generation via LLM
fruesome · reddit · 2026-08-16
MiDashengLM-Gen is an end-to-end framework using a pre-trained LLM and audio tokenizer. It employs per-token conditional flow matching for autoregressive, variable-length audio scene generation, blending speech, music, SFX, and acoustics coherently from text descriptions.
More from Multimodal
- InfinityEdit: Infinite Video Editing via Lightweight Adapter — Yunze Tong · 2026-08-24
- Seeking Audio Upscaling LLMs: Is There a 'Super-Resolution' Model for Music? — LeatherRub7248 · 2026-08-24
- Describe your dream world to an AI dragon, which generates the planet for you — repligate · 2026-08-24
- Using kintsugi texture to fix cracks in edited 3D meshes — repligate · 2026-08-24
- Generating Hannibal Character Videos with FL2VA Model — Nimblecloud13 · 2026-08-24
- MiniMax H3 Revives Medieval Short Stories: Complete Workflow Shared — zanatas · 2026-08-24