Treating Minimax H3 as a 'Dumb Cameraman' to Stitch a 2-Minute Dialogue Scene
niechta · reddit · 2026-08-09
A developer shared a complete workflow for creating a 2-minute dialogue scene using the Minimax H3 model. To avoid single-prompt failures, the author treated the AI like a "dumb camera operator," generating separate takes, coverage shots, and silent reaction handles, then manually editing them together.
Core Tech Stack & Parameters:
- Model & Plugins: Minimax H3 reference-driven path (Ref2VA), using ComfyUI 0.30 with ComfyUI-H3-Multishot and ComfyUI-GGUF nodes.
- Specs: 960×540 resolution, 24 fps; the model generates picture and 32 kHz stereo audio in a single pass.
- Inference: 20 steps, cfg 1.0, resmultistep / simple sampler, fixed seed.
- Hardware: 420 seconds per 362-frame generation on an RTX PRO 6000.
The author notes that the resolution limitation causes noticeable pixelation in cropped single shots, but post-processing successfully glued all chunks into a cohesive piece.
Related event: Developers Explore Diverse Workflows for MiniMax H3 Video Generation(6 posts)→
More from Multimodal
- MiniMax H3 R2V Model Test: Realistic Corridor Encounter Generation — BoothJudas9 · 2026-08-09
- Creating Dramatic Anime Videos Using MiniMax — No-Raspberry4782 · 2026-08-09
- Open Source ComfyUI Nodes: Fixing MiniMax H3 Fast Motion Artifacts — MatlowAI · 2026-08-09
- Beyond Prompting: The Case for an 'Aesthetics of Training' in AI Art — pixlpa · 2026-08-09
- MiniMax + Turbo LoRA Test: Impressive Music Generation and Audio Sync — LawfulnessRelevant45 · 2026-08-09
- MiniMax H3 Ultra Fast Tops HF Trending with Video Continuation and LoRA Support — realmrfakename · 2026-08-09