MiniMax H3 tutorial: Voice & Likeness LoRA training
shootthesound · reddit · 2026-08-19
A tutorial for training a single LoRA on MiniMax H3 for both voice and likeness using audio, video, and images. The demo LoRA was trained on 35 pics, 26 wavs, and a video clip. The guide includes a data prep tool called Gizmo. Strategy advice includes switching to audio-only training after epoch 40 to refine voice without overbaking visuals, noting that stills + audio is faster than video clips.
More from Multimodal
- Apple's Internalized Visual Thinking speeds up video reasoning — apple · 2026-08-19
- ComfyUI Cache Monitor Update: Manual Pinning and VRAM Freeing — Incognit0ErgoSum · 2026-08-19
- Dataset Release: 1M+ 19th-Century Public Domain Images with Masks — wjb_mattingly · 2026-08-19
- SciFigPlag-Bench: Benchmark Challenges LLMs on Figure Plagiarism — anshulkundaje · 2026-08-19
- Why Removing the Vision Encoder Can Be Better: An Infra Perspective — kastnerkyle · 2026-08-19
- LightOnOCR-2-1B: lightweight open-source OCR model tops 3M downloads on Hugging Face — adnan_hashmi · 2026-08-19