MiniMax Video Drift: Close-Ups Lose Identity After 3 Seconds
ItsMilaVoss · reddit · 2026-08-12
The author conducted in-depth tests on identity drift in MiniMax H3 image-to-video generation, finding that close-ups fall apart visually after 3 seconds, while over-the-shoulder shots can survive the full 6.5s.
Key Findings & Pitfalls
- Shot Size Dictates Usable Length: Close-ups (face filling the frame) drift into a different person (wider face, softer jawline) after 2.9s; medium shots 6.2s; wide shots last the full clip. Never cut a 6s close-up into two 3s pieces, as the second half is already someone else—generate separate 3s clips instead.
- Big Expressions Destroy Identity: Wide genuine laughs accelerate identity collapse. Keep expressions minimal and find energy in the edit.
- Faceless Shots Are Dangerous: Dark, undefined regions get filled with hallucinated people (even growing a second woman). Pack detail shots with actual objects.
- Measurable Prompts Work: "No push-in" is often ignored; use measurable constraints like "subject's head must occupy the same size in the first and last frame."
Additionally, small props (like necklaces) mutate silently and need explicit locking in prompts. Always use first-frame conditioning, never last-frame, to prevent the model from inventing a bad intro transition.
Related event: MiniMax H3 Video Generation Struggles with Character Consistency(3 posts)→
More from Multimodal
- Generating a Bollywood-Level Fight Scene for $0 Using Seedance 2.5 — taherdhanera · 2026-08-12
- Testing MiniMax H3 as a 3D Render Engine Inside ComfyUI — mickmumpitz · 2026-08-12
- Imperial College London's 4DGS VR Soldering Training Offers 10x Better Experience — OwariDa · 2026-08-12
- LTX 2.5 Faces the Classic SD3 Test with Impressive Generation Speed — WearNatural5992 · 2026-08-12
- AI tool turns selfie into campaign shot, no photographer or studio needed — SimplyAnnisa · 2026-08-12
- Using Narrative Logic to Guide AI Image Generation: Advanced Midjourney Prompting — tisch_eins · 2026-08-12