MiniMax H3 workflow: REF2VA visuals + FL2VA audio + LightX2V speed, 198s on RTX 5090

gabxav · reddit · 2026-09-05

After testing MiniMax H3 since release, gabxav found REF2VA and FL2VA each excel at different things and shares a ComfyUI workflow combining their strengths.

Findings

The workflow

All reproduction assets are shared: H3 models and LoRA on Hugging Face, workflow JSON, voice/image reference files, KJNodes and AudioRefine. Reference performance: 198.4 seconds per generation on an RTX 5090.

Original post →

More from Multimodal

Multimodal channel →