Native ComfyUI Workflow Does Instant Character/Style References With Voice, No refmod Needed
crinklypaper · reddit · 2026-09-12
A Redditor shared a near-native ComfyUI workflow that replicates refmod-style instant character and style reference videos — with voices — using no refmod or custom exotic nodes, with the workflow file and breakdown included.
- Images from a folder are turned into video frames and fed as a reference video (a grid version packs images into a single reference image; the video version is preferred).
- Audio: feed a single mp4 via VHS node, or truncate the first seconds of 4 clips and concatenate into a clean 15–30s voice clip.
- Built on KJ nodes, native nodes, and VHS; swap datasets by pointing to a folder — no safetensors management or audio editing.
- Prompts follow MiniMax-H3's ref syntax (an LLM can write them); add more description if dataset traits bleed into generations.
- Known caveat: a ComfyUI memory-management bug means clearing cache or restarting after dataset changes; video-clip inputs are untested.
More from Multimodal
- GPT Image 2.5 lands in CapCut PC, powering a 3-step idea-to-film AI workflow — thetripathi58 · 2026-09-12
- GPT Image 2.5 is coming to CapCut PC via Design Studio and AI Image — thetripathi58 · 2026-09-12
- Pixel art to playable game: a full pipeline demo via Sprite Fusion API — HugoDuprez · 2026-09-12
- AI's Take on 'Grandma Games' Produces Delightfully Weird AI Video Concept — Aiden_Tech_Ai · 2026-09-12
- TesanaAI launches image-to-game: turn any screenshot into a playable game in under 5 minutes — gabriel1 · 2026-09-12
- Suno can now train on copyrighted music, user demos a track — AndyMasley · 2026-09-12