ScenA on Hugging Face: Generates Two-Speaker Dialogue and Scene Sound Effects in One Pass

multimodalart · x · 2026-08-08

A new omni-TTS model named ScenA has been released on Hugging Face, described as "criminally under-hyped."

Fine-tuned from the LTX audio generation module, the model takes a text description and two voice references to produce a fully mixed scene in a single pass. It can generate dialogue between two speakers and automatically includes relevant scene sound effects.

Original post →

More from Multimodal

Multimodal channel →