ContextMaster Unifies Multi-Shot Video Creation at 16 FPS on a Single GPU
KlingTeam · hf · 2026-08-07
ContextMaster formalizes a new setting called Interactive Multi-Shot Video Creation (IMVC), enabling a single model to maintain a shared history across generation, reference conditioning, and editing operations.
To prevent context read costs from scaling with history length, the model introduces fixed-budget sparse context routing combined with ConstraintSink to keep task constraints visible. It also proposes a two-stage privileged context distillation framework: transferring full-context behavior from a dense teacher via consistency distillation, followed by deployment refinement using distribution matching.
Experiments show superior task fulfillment and cross-shot consistency over specialized baselines, while achieving 16 FPS inference on a single GPU.
More from Multimodal
- MiniMax H3 Test: Creating a Humorous and Dynamic Music Video — uhf789 · 2026-08-07
- ByteDance Seedance 2.5 Test: 49s Video with 4 Coherent Characters — SucceededMind · 2026-08-07
- Generating 30s Videos from a Single Image: MiniMax H3 Local Test & Prompt Breakdown — Tight_Organization54 · 2026-08-07
- Seedance 2.5 Supports Up to 30 Reference Inputs — OdinLovis · 2026-08-07
- MiniMax H3 Test: Stunning Visuals, But Still Fails at Door Logic — martinerous · 2026-08-07
- Seedance 2.5 Tested: Native 30s 4K Output and Multi-Reference Consistency — SucceededMind · 2026-08-07