MiniMax H3 pure text-to-video test with a full shot-by-shot prompt
BitterAd8431 · reddit · 2026-08-28
A frequent MiniMax H3 user tried prompt-only text-to-video (usually they use reference images), with an integrated LLM helping expand the prompt, producing a comedic manga-style clip of a white-haired cat-eared girl triggering a chain of kitchen fires.
The full prompt is shared, with a reusable structure:
- Shot-by-shot description covering character appearance, actions, scene changes and plot beats;
- Second-precise timeline control: e.g., pure nonverbal action from 0–2.88s, dialogue from 2.88–9.88s, reactions/ambience filling 9.88–14.38s;
- Audio-visual constraints: no human voices, breathing, or mouth movement outside the tagged dialogue window, with continuous ambience and synced SFX.
More from Multimodal
- Omni, Qwen, and other Flash models tested; video generated under 10s — iScienceLuvr · 2026-08-28
- Midjourney v8.2 Editing Features Tested: Removal & Repaint — LudovicCreator · 2026-08-28
- Krea AI launches new Krea Three model at NYC event — chrisfirst · 2026-08-28
- Google DeepMind Releases ORBIT++ Benchmark for SfM in 360° Video — taiyasaki · 2026-08-28
- Peking Univ & Kling Unveil MAVIN to Solve Multi-Shot Video Narrative Challenges — 机器之心 · 2026-08-28
- Midjourney tests V8.2 edit model with inpainting, outpainting, and image-to-image — DavidSHolz · 2026-08-28