2:47 Mini-Documentary Fully Local on 16GB: MiniMax H3 Workflow Breakdown

Short_Regular_7191 · reddit · 2026-08-10

From a 5-Second Clip to a 3-Minute Documentary

A creator shared the complete workflow for generating a 2 min 47 sec mini-documentary about the Trial of Socrates entirely locally on a single RTX 5060 Ti 16GB using MiniMax H3. The film consists of 36 clips across 4 scenes, maintaining one consistent character throughout with exact lip-sync and invisible joins.

Key Engineering Breakthroughs

Instead of crossfades, the last frame of clip N is passed as the first frame for clip N+1. The model picks up exactly from there (SSIM 0.89), making cuts indistinguishable from natural motion on the timeline.

A full stop inside the <d> dialogue tag equals a 1-second dramatic pause. Use commas for tight timing. Never prompt that a sentence "gets cut off," as the model will invent words to stretch it.

The reference face tends to appear on extras. Mitigations: don't re-declare fullypreserved in continuation clips, explicitly state "only one man has this face," and run an InsightFace QA pass on every face.

The take selector needs optical flow, net camera displacement, and trajectory straightness to detect real camera movement vs. jitter. When face-similarity and sharpness conflict, a human must decide.

Video is trimmed to continuous narrator audio. Traps: concatenating AAC via stream-copy accumulates drift (extract per-segment PCM first); AAC encoder adds 0.3 dB (measure true peak); Whisper medium silently fixes TTS grammatical errors (use large-v3 for QA).

Production Numbers

Related event: Local MiniMax H3 Video Generation Workflows Tested on 16GB GPUs(5 posts)→

Original post →

More from Multimodal

Multimodal channel →