Turn Any Local LLM Into a MiniMax H3 Video Prompt Assistant
Due-Quiet572 · reddit · 2026-08-05
A developer shared a system prompt that transforms local vision-language models (like Qwen3-VL-30B) into a dedicated prompt assistant for MiniMax H3 video generation.
Workflow:
- Users initiate the assistant, which then asks step-by-step questions about language, duration, aspect ratio, camera movement, and audio details, rather than dumping a long form.
- Supports multiple input modes: text-only, first/last frame images, multiple reference images (assigning roles like character, location, clothing), and reference videos.
- Automatically generates the final prompt in English formatted specifically for MiniMax.
Configuration Tips:
- Recommends using LM Studio to load the GGUF quantized model alongside its vision projection (.mmproj) file.
- Advises increasing context length (32768) and max output tokens (8192) to prevent truncation of longer prompts.
Related event: Developers Create Prompt Assistant for MiniMax H3 Video Model(2 posts)→
More from Multimodal
- Text-to-Video vs Image-to-Video: Impressive Demos vs Practical Usability — EntireBig7258 · 2026-08-06
- Treblo launches open-source AI music classifier with <0.1% false positive rate — The Verge AI · 2026-08-06
- MiniMax T2V accurately generates Adam Sandler and Will Ferrell with zero references — RainbowUnicorns · 2026-08-06
- Running MiniMax H3 on Dual RTX 5060 Ti: VRAM Splitting and Unloading Strategies — Kahvana · 2026-08-06
- Oscar-Nominated Cinematographer to Judge First-Gen AI Filmmakers at Higgsfield Festival — SimplyAnnisa · 2026-08-06
- MiniMax H3 Test: Generates Video with Sound Effects in 120s on RTX 5090 — R34vspec · 2026-08-06