MiDashengLM-Gen: Unified Audio Scene Generation via LLM

fruesome · reddit · 2026-08-16

MiDashengLM-Gen is an end-to-end framework using a pre-trained LLM and audio tokenizer. It employs per-token conditional flow matching for autoregressive, variable-length audio scene generation, blending speech, music, SFX, and acoustics coherently from text descriptions.

Original post →

More from Multimodal

Multimodal channel →