Training an audio-drama LoRA: which architecture or platform?

wh33t · reddit · 2026-09-22

A Reddit user asks how one would train an "audio drama" generator: TTS, music, video, and sound-effect generators all exist separately, but nothing rolls dialogue plus ambient effects (creaking doors, bootsteps, tiger roars) into a single promptable system. They ask what architecture or platform would suit training such a LoRA to get "90% of the way there" from a script. Open discussion question, no answers or implementation details in the post.

Original post →

More from Multimodal

Multimodal channel →