Can MiniMax H3's R2V Capability Be Used for Reference-to-Image?

ok-onwrap · reddit · 2026-08-19

A user is exploring how to leverage MiniMax H3's "reference-to-video" (R2V) capability for single-frame "reference-to-image" (R2I) generation. H3 appears to have a minimum output of 5 frames, and the user seeks methods to force it to generate just one frame, potentially via parameters or ComfyUI workflows.

The goal is to find a free or open-source model that matches H3's ability to accept multiple reference images for consistent character/scene generation, similar to GPT Image 2 but without the cost. The user is asking for experiments or workflows that achieve this.

Original post →

More from Multimodal

Multimodal channel →