MiniMax H3 workflow: REF2VA visuals + FL2VA audio + LightX2V speed, 198s on RTX 5090
gabxav · reddit · 2026-09-05
After testing MiniMax H3 since release, gabxav found REF2VA and FL2VA each excel at different things and shares a ComfyUI workflow combining their strengths.
Findings
- REF2VA: noticeably better visuals — natural skin texture, lighting, environments.
- FL2VA: over-smooth "AI look," but handles high-motion characters better, and its voice/audio cloning is significantly cleaner (less echo, noise, artifacts).
The workflow
- REF2VA for main video generation (quality)
- FL2VA for audio refinement
- LightX2V 8-step LoRA for a major speed boost
- H3 AudioRefine node to pass generated audio through FL2VA for cleaner sound
All reproduction assets are shared: H3 models and LoRA on Hugging Face, workflow JSON, voice/image reference files, KJNodes and AudioRefine. Reference performance: 198.4 seconds per generation on an RTX 5090.
More from Multimodal
- Seedance 2.5 nails emotional progression and believable aerial physics in demo — azed_ai · 2026-09-05
- H3 zombie apocalypse short: AI video hits surprisingly hard emotionally — ArjanDoge · 2026-09-05
- Fine-tuning Meta SAM 3 for narrative-aware copyspace: fixing AI picture book typography — andrew_n_carr · 2026-09-05
- Blender clay renders + video diffusion: a workflow idea for OpenAI's Astra agent — OdinLovis · 2026-09-05
- Reddit user compares two MiniMax H3 Turbo LoRAs for video generation — No-Bee-231 · 2026-09-05
- AI short film "Ishiro - The Last Blade" shared on Reddit — Nervous-Barnacle-888 · 2026-09-05