Scenema Audio Hits ComfyUI: Expressive TTS & Zero-Shot Cloning on 8GB VRAM
a__side_of_fries · reddit · 2026-08-06
Scenema Audio is now available as a native ComfyUI custom node. The model has been quantized to run efficiently on as little as 8GB VRAM (e.g., RTX 3070), with generation speeds reaching up to 2x realtime.
Key Features:
- Features zero-shot voice cloning. Users can describe specific emotions (rage, grief) and use inline bracket cues (like [he laughs softly]) to trigger actions or tone shifts at exact moments.
- Ships with 12 preset voices covering various accents, ages, and emotional registers.
Requirements & Limitations:
- The first run downloads about 30GB of weights. The text encoder uses Gemma 3 12B, requiring users to accept its HuggingFace license and set up an HFTOKEN.
- As a diffusion model, some seeds may produce repetition or gibberish, making it best suited for a post-editing workflow. Phonetic spelling is recommended for proper nouns.
Related event: Scenema Audio Runs on 8GB VRAM via ComfyUI(2 posts)→
More from Multimodal
- Open-Source Project Clones Zack D. Films 3D Animated Short Workflow — matchaman11 · 2026-08-06
- Using Claude with Local ComfyUI MCP for Conversational Video Generation — EndPsychological8822 · 2026-08-06
- Is ControlNet obsolete? Krea 2-Pose showcases precise image generation control — multimodalart · 2026-08-06
- AI Video Generation Test: Recreating a Classic Sailor Moon Scene — VinceTrust · 2026-08-06
- Minimax H3 video generation stuck in 'uncanny valley', dev says — mattshumer_ · 2026-08-06
- Grok Imagine to Introduce Keyframe Support and Upgraded Image Model — chaitu · 2026-08-06