Running MiniMax H3 video generation locally on a 12GB laptop: face detail is the bottleneck

sarasa_0505 · reddit · 2026-08-23

A user runs MiniMax H3 (416P) locally in ComfyUI on an RTX 4000 Laptop GPU (12GB VRAM) + 32GB RAM. After getting 416P working, they hit bottlenecks: 480P causes OOM; faces lose detail and distort in wide shots; 4x-UltraSharp upscaling sharpens overall but exaggerates face distortion; and they want more natural Japanese speech and lip-sync. A full ComfyUI workflow screenshot is attached, with requests for face-restore nodes, tiled upscalers, VRAM optimization, and lip-sync node recommendations.

Original post →

More from Multimodal

Multimodal channel →