Local Fine-Tuning MiniMax H3 Video Model Tested on a Single RTX 5090
No_Statement_7481 · reddit · 2026-08-07
A developer shared their practice of locally fine-tuning the MiniMax H3 video model using the Osris Ai toolkit. They built a dataset of 1-second, 512x512 resolution videos for LoRA training tests.
Hardware & Performance:
- With Layer offloading enabled on a single RTX 5090, VRAM usage was extremely low (6.6GB / 32GB VRAM), but consumed 78GB of system RAM.
- Each iteration step took about 3.43 seconds, meaning 3000 steps will take roughly 3 hours.
Pitfalls & Tuning:
- Frame count is crucial: failing to adjust the default 39 frames to the dataset's 22 frames resulted in samples with "chipmunk sounding" audio.
- The author currently used simple captions and plans to test the limits with a better dataset and proper prompting standards next.
Related event: MiniMax H3 Video Model Supports Local LoRA Training on 16GB VRAM(3 posts)→
More from Multimodal
- User showcases beautiful world generated by LTX 2.5 model — cocktailpeanut · 2026-08-15
- Don't Overlook Minimax's Video Editing: Comparable to Google Omni — Radyschen · 2026-08-15
- First Decent Generation Using Minimax-H3 on RTX 3060 12GB — solomars3 · 2026-08-15
- Pika Launches Audio API: Sound Effects and Speech from $0.0002/sec — minchoi · 2026-08-15
- MiniMax H3 video generation test: 8-step Turbo LoRA impressive — serap98765 · 2026-08-15
- MiniMax H3 video demo: Generating medieval scenes in 5-8 seconds — Queasy-Breakfast-949 · 2026-08-15