Realtime Interactive Video on MiniMax H3: Character LoRA Training Puzzles and a Key ComfyUI Conversion Trap

Big-Set9728 · reddit · 2026-09-07

A developer built a realtime interactive video scene on MiniMax H3 — you type lines and the character answers with generated audio, with 5s clips generating slightly faster than realtime. The bottleneck is character consistency: they want a trained character LoRA instead of paying for reference conditioning each generation.

Key findings and a trap:

Open questions: whether a base-H3 character LoRA transfers to the distilled checkpoint, whether it can stack with the distill LoRA, if T2I stills suffice for talking-face identity, and multi-GPU sequence parallelism numbers. Author offers paid consulting.

Related event: Dev Tests Real-Time Interactive Video with MiniMax H3(2 posts)→

Original post →

More from Multimodal

Multimodal channel →