Scenema Audio Launches Native ComfyUI Node, Runs on 8GB VRAM
a__side_of_fries · reddit · 2026-08-06
Scenema Audio has released a native ComfyUI custom node. By quantizing the model, it now runs on as little as 8GB VRAM, solving the previous issue of full-precision transformers being too heavy for self-hosting.
Key Features:
- Expressive text-to-speech (TTS) with zero-shot voice cloning.
- Allows text-based emotion descriptions (e.g., rage, grief) and inline stage directions (like [he laughs softly]), which the model performs precisely at the exact spot.
- Ships with 12 preset voices covering various accents, ages, and emotional registers.
Deployment & Limitations:
- Requires a one-time download of 30GB weights. The text encoder uses Gemma 3 12B (requires accepting HuggingFace license and configuring a token).
- As a diffusion model, some seeds may produce repetition or gibberish. It is designed for a post-editing workflow: generate, pick the best take, and trim.
Related event: Scenema Audio Runs on 8GB VRAM via ComfyUI(2 posts)→
More from Multimodal
- Community Calls for Quantized MiniMax H3 for Apple Silicon — sockenull · 2026-08-06
- Hailuo H3 Meets Adobe AE: New Plugin Enables In-App Video Generation and Object Replacement — aziz4ai · 2026-08-06
- Developer Builds Local Multimodal AI to Generate Personalized Audiovisual Works — Merzmensch · 2026-08-06
- Video Generation Showdown: Seedance 2.5 vs. Minimax H3 — LudovicCreator · 2026-08-06
- Turning a Single Product Image into an AI Commercial: Full Workflow — Gullible-Goose-1992 · 2026-08-06
- 'I don't want to play with you anymore': Hilarious Text-to-Video Fail — Superb-Painter3302 · 2026-08-06