MiniMax H3 Hybrid Model: Merging First Frame and Reference Images
Tokey_TheBear · reddit · 2026-08-17
A user shared a hybrid model approach for MiniMax H3 (merging FL2VA and REF2VA) to solve the limitation where official models cannot simultaneously maintain high-quality first-frame locking and use extra reference images.
Key Solution:
- Uses fl2va as the base for quality while incorporating transformer blocks from ref2va to enable reference conditioning.
- The b20-49 variant is preferred for high reference faithfulness (e.g., glyphs, specific stills), while b30-49 offers slightly higher raw quality.
Use Cases:
- Pattern A (First frame + Glyph Ref): Locks the opening scene location while using a second image (e.g., a logo) purely as a style/symbol reference without turning it into a video frame.
- Pattern B (Timed Hard Cuts): Enables precise hard cuts to different still images at 0s, 3s, and 6s within a single clip, overcoming the stock FL2VA limitation of only supporting start and end frames.
Implementation requires specific prompt formatting (e.g., fullypreserved) and ComfyUI wiring using the Reference-to-Video workflow.
More from Multimodal
- Does Minimax H3 Require Strict Native Resolution of 1344x768? — rm_rf_all_files · 2026-08-17
- Seeking Model to Optimize Minimax H3 Image-to-Video Prompts — waseem335 · 2026-08-17
- How to Successfully Convert Style with Minimax H3 Ref2V? — Adventurous-Gold6413 · 2026-08-17
- Zero assets modeled: OPUS 5 one-shots a 1999 arcade-style game — RanaHanocka · 2026-08-17
- AI news broadcast short film made with Flux and H3 reference model — JuniorEnsign · 2026-08-17
- Minimax H3 local test: 6-sec video takes 38 min on 4070 Ti Super — Azhram · 2026-08-17