ScenA on Hugging Face: Generates Two-Speaker Dialogue and Scene Sound Effects in One Pass
multimodalart · x · 2026-08-08
A new omni-TTS model named ScenA has been released on Hugging Face, described as "criminally under-hyped."
Fine-tuned from the LTX audio generation module, the model takes a text description and two voice references to produce a fully mixed scene in a single pass. It can generate dialogue between two speakers and automatically includes relevant scene sound effects.
More from Multimodal
- Minimax Video Test: Creating Coherent Fight Scenes with Krea — BigDovahkiin · 2026-08-08
- Claude Opus Generates Trippy ASCII Art Video Pulsing to Music in Stunning Demo — repligate · 2026-08-08
- Pure text-to-video generates cinematic Captain Jack Sparrow with stunning accuracy — Only_Voice569 · 2026-08-08
- Midjourney + Seedance 2.5: AI Short Film Workflow for a Forest Walk — miilesus · 2026-08-08
- Suno tightens AI music generation rules to fight spam and copyright issues — The Decoder · 2026-08-08
- Mom Loves Korean AI Hip Hop: 'It's So Over Family' — yungcontent · 2026-08-08