MiniMax character LoRA on Runpod: 150 images, ~2300 steps to nail likeness, and 5090 beats H100
Draufgaenger · reddit · 2026-09-04
A practitioner shares hands-on findings from training MiniMax character LoRAs on Runpod (with a 2-minute video tutorial):
- Datasets: image-only datasets worked much better than mixed ones; focus on near-frame-filling, sharp faces; MiniMax needs more images than LTX/Wan (150 images for a blonde woman's likeness, 50 sufficed for the author's own face);
- GPU surprise: RTX 5090 beat the H100 on this image-only workload; RTX 6000 PRO was only 20% faster than the 5090;
- Step counts: the author's face peaked at 700 steps; the woman's dataset peaked at 2310 steps — and notably dataset size didn't move the peak (2300 steps regardless of 50 or 150 images), though the 150-image LoRA was objectively much better;
- Data noise matters: 4 brown-haired images out of 150 caused brown-haired outputs past the peak even with blonde prompts;
- Natural-prose captions including lighting and look work best; voice training isn't supported in diffusion pipe, but MiniMax accepts a reference voice.
More from Multimodal
- User edits cinematic Tesla Cybercab video entirely with Grok Build — elonmusk · 2026-09-04
- Omni 1.1 video references impress: turning a car into a boat while staying faithful — fofrAI · 2026-09-04
- Redditor fine-tunes SDXL on 60 childhood photos to simulate memory recall — uisato · 2026-09-04
- Recreating a viral fight video with MiniMax H3: full reference-generation workflow — MixZealousideal9359 · 2026-09-04
- Midjourney one-word prompting: obscure dialect word 'Sillion' makes a striking image — tisch_eins · 2026-09-04
- Prompt-to-World: Claude Plus Thrixel's Build World Skill Spawns Interactive 3D Worlds — RanaHanocka · 2026-09-04