OmniChar adds consistent voice cloning: 10-second sample locks face, body and voice
ashishsanu · reddit · 2026-10-06
A developer extended the OmniChar portable character format (.char) on the MiniMax H3 + ComfyUI workflow: after achieving consistent face, clothes and body across videos, it now supports consistent voice — drop a 10–30 second mp3/wav sample and the character speaks in the cloned voice with lip sync (provided natively by H3).
- Updated ComfyUI node (nightly), with workflows for encoding a .char with voice, MiniMax H3 voice generation, and full face/body/cloth consistency
- Limitations: strong in English, may babble in other languages; avoid multiple speakers in the sample
- GPLv3 open source, supports krea2, character finetuning, and a community character library
Related event: OmniChar Adds Voice Consistency to .char Character Format(2 posts)→
More from Multimodal
- Qwen Image 2.1 Uncensored MCP: a full image pipeline in 12GB VRAM — Delicious-Farmer-234 · 2026-10-06
- 8 prompts turn Claude Opus 5.5 in Claude Code into an explainer video studio — alex_verem · 2026-10-06
- MetaCanvas lands NeurIPS: MLLMs plan directly in diffusion latent space — mohitban47 · 2026-10-06
- Ming Image 0.1 in ComfyUI: free AI model tested for text-heavy design work — pixaromadesign · 2026-10-06
- One-word Midjourney v8.2 prompt: 'Diaphanous' renders sheer translucency — tisch_eins · 2026-10-06
- Reddit User Shares AI-Reimagined 'Princess and Dragon' Short Video — Throwaway350750 · 2026-10-06