MiniMax H3 tutorial: Voice & Likeness LoRA training

shootthesound · reddit · 2026-08-19

A tutorial for training a single LoRA on MiniMax H3 for both voice and likeness using audio, video, and images. The demo LoRA was trained on 35 pics, 26 wavs, and a video clip. The guide includes a data prep tool called Gizmo. Strategy advice includes switching to audio-only training after epoch 40 to refine voice without overbaking visuals, noting that stills + audio is faster than video clips.

Original post →

More from Multimodal

Multimodal channel →