Realtime Interactive Video on MiniMax H3: Character LoRA Training Puzzles and a Key ComfyUI Conversion Trap
Big-Set9728 · reddit · 2026-09-07
A developer built a realtime interactive video scene on MiniMax H3 — you type lines and the character answers with generated audio, with 5s clips generating slightly faster than realtime. The bottleneck is character consistency: they want a trained character LoRA instead of paying for reference conditioning each generation.
Key findings and a trap:
- FastVideo's FastH3 (4-step DMD2 distill) does 6s for a 5s clip on 4x B200 and ships a pre-extracted LoRA with lorapath/lorastrength hooks
- diffusion-pipe and ai-toolkit both list H3 support
- VSA weights don't survive stock ComfyUI conversion: 50 togatecompress tensors are silently dropped, producing noise; dense converts cleanly
Open questions: whether a base-H3 character LoRA transfers to the distilled checkpoint, whether it can stack with the distill LoRA, if T2I stills suffice for talking-face identity, and multi-GPU sequence parallelism numbers. Author offers paid consulting.
Related event: Dev Tests Real-Time Interactive Video with MiniMax H3(2 posts)→
More from Multimodal
- Open-sourced ComfyUI nodes bring Reactor.inc hosted video models including MiniMax FastH3 — TheMoonMidas · 2026-09-07
- GPT-6 'Astra' tested: one image auto-modeled into a 3D Blender city, unverified — gabrielchua · 2026-09-07
- Touching up old Midjourney images with GPT shows how far image models have come — mimi10v3 · 2026-09-07
- Astra builds full gothic cemetery 3D assets in Blender in 28 minutes — nptacek · 2026-09-07
- AI agent runs 12-hour experiments to autonomously build a fighting game — msg · 2026-09-07
- He spent $3,000 using Seedance to make his own mecha film — Th3dzon33 · 2026-09-07