Experimental ComfyUI Nodes: Adapting MiniMax H3 Video Model for Image Generation

killerciao · reddit · 2026-08-03

A developer has created a custom ComfyUI extension that adapts the MiniMax H3 video model for text-to-image, image-to-image, and reference editing.

Instead of forcing the model to generate a single frame—which typically yields poor results—the workflow generates a short temporal sequence, decodes the minimum required frame packet, selects the best still, and outputs only that image. The author notes that while the method works decently for image editing, H3 is fundamentally a video model, meaning softness, blockiness, banding, and grid artifacts can still be present.

Original post →

More from Multimodal

Multimodal channel →