MovieGrid uses multi-grid post-training to make long-form multi-shot video generation coherent

Jiawei Mao · hf · 2026-09-09

A new method called MovieGrid decomposes long videos into spatially arranged chunks for joint modeling, improving multi-shot coherence while scaling video length efficiently. Published on Hugging Face by Jiawei Mao, it targets long-form multi-shot video generation via multi-grid post-training.

Original post →

More from Multimodal

Multimodal channel →