ContextMaster Unifies Multi-Shot Video Creation at 16 FPS on a Single GPU

KlingTeam · hf · 2026-08-07

ContextMaster formalizes a new setting called Interactive Multi-Shot Video Creation (IMVC), enabling a single model to maintain a shared history across generation, reference conditioning, and editing operations.

To prevent context read costs from scaling with history length, the model introduces fixed-budget sparse context routing combined with ConstraintSink to keep task constraints visible. It also proposes a two-stage privileged context distillation framework: transferring full-context behavior from a dense teacher via consistency distillation, followed by deployment refinement using distribution matching.

Experiments show superior task fulfillment and cross-shot consistency over specialized baselines, while achieving 16 FPS inference on a single GPU.

Original post →

More from Multimodal

Multimodal channel →