Musubi Tuner merges MiniMax H3 video-audio model training support
bdsqlsz · x · 2026-09-16
kohya announced that Musubi Tuner has merged support for MiniMax-H3 (text-to-video-with-audio) training into main, with base implementation contributed by sdbds.
- The official guide (minimaxh3.md) covers what to download, choosing a training recipe, and commands for caching, training, sampling, and generation
- Two companion docs: minimaxh31f.md for one-frame (image) generation/training (time indices, editing/inbetween datasets, reference-conditioned images) and minimaxh3advanced.md on timestep/loss internals, guidance loss, teacher matching, caching, quantization, and generation internals
- The author notes the subject-reference teacher learning dataset setup is confusing and will document it separately
More from Multimodal
- Full workflow breakdown: making a 6-minute AI duet with Grok and Suno — Kyrannio · 2026-09-16
- 3D modeler tests turning a friend into a figure from just photos with AI — Popular_Double4000 · 2026-09-16
- This AI-made film cost $29,575 and took just 10 days to make: Nexus Ep 4 by PJ Ace — NightsRadiant · 2026-09-16
- Unfold wins two awards at OpenAI Singapore hackathon for turning 2D manuals into 3D guides — cedric_chee · 2026-09-16
- Unverified GPT-6 Astra demo rebuilds indoor scenes in 3D from one reference image — Dazzling_Gate_330 · 2026-09-16
- Creator recreates viral volcano popcorn meme entirely with flovaai — SimplyAnnisa · 2026-09-16