Xiaomi releases MiDashengLM-Gen for one-prompt scene audio generation

solyarisoftware · x · 2026-08-16

Xiaomi released MiDashengLM-Gen on Hugging Face. It generates six tagged views of a scene (caption, transcript, voice, music, sfx, room) from a single prompt, rendering the whole scene as a single 16kHz audio clip. A demo is available on Spaces.

Original post →

More from Multimodal

Multimodal channel →