MiniMax H3 in Practice: Native Audio-Video Generation and Advanced Prompting
Smyshnikof · reddit · 2026-08-06
A developer shared their hands-on experience and prompting engineering details after porting the GalaxyAce LoRA to the MiniMax H3 model.
Model Features:
- H3's standout feature is generating video and audio natively in a single pass (room tone, traffic, speaking), similar to the Sora 2 experience, with open weights.
- The 33B model ships prequantized at around 43GB, requiring massive VRAM. Running it realistically requires an RTX 5090 (32GB) or rented Blackwell GPUs.
Prompting Techniques:
- Effective Structure: Scene => Beats with timecodes => Camera behavior => Audio => No text/logos.
- Pitfalls: Avoid words like "cinematic," "film grain," or "depth of field" (which imply expensive lenses), as the model tries to simulate a cheap camera sensor.
- Provided a full copy-paste prompt example, noting that spoken dialogue can easily be swapped into other languages.
Training Details: Used Ostris AI-Toolkit, rank 32, taking roughly 1 hour.
More from Multimodal
- Fully AI-Generated Sci-Fi Series: Runway Powers Complete Pilot Episode — bennash · 2026-08-06
- MiniMax H3 Test: Generating Absurd Chase Sequence on a 4090 in 10 Minutes — Metapharstic · 2026-08-06
- Testing MiniMax for Image Generation: Character Replacement and Outpainting — Sudden_List_2693 · 2026-08-06
- Analyzing Coherent Mutating Videos: AI Vid2vid or Datamoshing? — showerelf · 2026-08-06
- AI Video Generation: Tom and Zendaya's 90s South Indian Wedding — njan_ninde_thanda · 2026-08-06
- AI Generated Meme: Bart Simpson as a Hyperstitional Object — fuzzhello · 2026-08-06