Fizgig 4.0: train Minimax H3 with photos, audio and video in one dataset

shootthesound · reddit · 2026-08-18

Open-source training tool Fizgig 4.0 is out:

The author's key hands-on takeaway: with the right settings, H3 handles image-based training without killing its video ability; combining photos and wavs is super fast and makes voice training easy (a shared trigger word is recommended). Video training works but is unavoidably slower — if you're not teaching anything beyond what photos and audio can convey, skip video; use it when you need to capture motion. A detailed YouTube tutorial follows.

Original post →

More from Multimodal

Multimodal channel →