Fizgig 4.0: train Minimax H3 with photos, audio and video in one dataset
shootthesound · reddit · 2026-08-18
Open-source training tool Fizgig 4.0 is out:
- Combine Minimax H3 video, audio (wav/mp3) and photo training into a single dataset
- New Gizmo dataset prep tool for easy video/audio dataset preparation
- Turbo mode, plus a big INT8 speedup for 16GB-VRAM users (community contribution from rintic-13)
The author's key hands-on takeaway: with the right settings, H3 handles image-based training without killing its video ability; combining photos and wavs is super fast and makes voice training easy (a shared trigger word is recommended). Video training works but is unavoidably slower — if you're not teaching anything beyond what photos and audio can convey, skip video; use it when you need to capture motion. A detailed YouTube tutorial follows.
More from Multimodal
- AI Short Film 'DO NOT WAKE HER' Showcases Visual Effects — RealisticValuable484 · 2026-08-18
- MiniMax Character Swap Test: Person-to-Person and Animal-to-Animal Work — Alex-edits123 · 2026-08-18
- Qwen-Video-Edit repurposes an image editing DiT to edit videos, no video-pretrained backbone needed — qixing_huang · 2026-08-18
- Minimax H3 Generates a 'Deleted Scene' From the Doomsday Trailer — beatlepol · 2026-08-18
- Stabilizing renders with depth maps and object tracking — alecubudulecu · 2026-08-18
- Minimax H3 Ignores Image Input in ComfyUI Workflow, User Reports — magik_koopa990 · 2026-08-18