Running Online RL on MiniMax-H3 Joint Audio-Video Generation with DiffusionNFT
青稞AI · wechat · 2026-09-08
The VeRL-Omni team documents running a complete online RL loop for MiniMax-H3 — an open-weights joint audio-video generation model (up to 15s, 2K, 24FPS, stereo audio; Elo 1227, top-3 on Text-to-Video Arena and the only open-weights model there).
- Using DiffusionNFT, they close the loop for T2VA (text-to-audio-video) and FL2VA (first/last-frame-conditioned) training: vLLM-Omni handles rollout, Diffusers/FSDP2 trains the actor, and VeRL-Omni wires data, rewards, NFT loss, and LoRA old-policy syncing together.
- Rewards combine CLAP (text-audio alignment) and ImageBind (audio-video consistency) to close optimization gaps; key pitfalls include H3's CFG distillation, data-fraction timesteps, and inverted velocity sign, where silent mismatches can hide behind healthy-looking loss curves.
- On 8 GPUs, core reward mean rose from 0.41 to 0.51 in 3+ steps with stable gradnorm (0.06–0.12); a full reproducible config is provided.
More from Multimodal
- Reddit user shares WIP AI-generated dark fantasy short film 'Wanderers' — DaWid_Shapiro · 2026-09-11
- Creator makes 2D electro-pop anime music video with just a prompt using MiniMax H3 — Hailuo_AI · 2026-09-11
- Street View to driving footage: GPT Astra fetches images, MiniMax H3 turns them into dashcam video — Hailuo_AI · 2026-09-11
- Single-author ECCV 2026 paper makes rolling shutter correction practical — ducha_aiki · 2026-09-11
- AI digital human covers Japanese classic so realistically viewers can't tell — JourneymanChina · 2026-09-11
- ComfyUI Style Explorer Adds LoRA Preview Catalog and Sharing — neonsparksuk · 2026-09-11