Training an audio-drama LoRA: which architecture or platform?
wh33t · reddit · 2026-09-22
A Reddit user asks how one would train an "audio drama" generator: TTS, music, video, and sound-effect generators all exist separately, but nothing rolls dialogue plus ambient effects (creaking doors, bootsteps, tiger roars) into a single promptable system. They ask what architecture or platform would suit training such a LoRA to get "90% of the way there" from a script. Open discussion question, no answers or implementation details in the post.
More from Multimodal
- AI Short Film 'Sleep Paralysis Demon' Made with MiniMax H3 + SeedVR2 — Warp_d · 2026-09-22
- Face Swap with Qwen Image 2.1: No Masking, Just Two Images and a Prompt — sci032 · 2026-09-22
- MiniMax H3 Max Tops Design Arena Video Editing Leaderboard with Elo 1373 — noahsolomon · 2026-09-22
- Redditor Turns Home-Shot Clip into an AI Action Scene (Before & After) — nav132 · 2026-09-22
- Nano Banana Pro crushes GPT Images 2.5 in same spaceship prompt test — opmgyhx · 2026-09-22
- Grok 4.7 jumps to #3 on BuildingBench 3D generation, 66% cheaper than Fable 5.1 — ZhitingHu · 2026-09-22