ShotPlan adds learnable planning tokens for cinematic multi-shot video generation

Tele-AI · hf · 2026-07-21

What the paper proposes

ShotPlan targets cinematic video generation, where single-shot generation is not enough and coherent multi-shot structure matters.

Reported outcome

The authors say ShotPlan outperforms existing cinematic video generation methods, with:

In short, the work is about making video generation behave more like a directed sequence of shots rather than a single uninterrupted clip.

Original post →

More from Multimodal

Multimodal channel →