ShotPlan adds learnable planning tokens for cinematic multi-shot video generation
Tele-AI · hf · 2026-07-21
What the paper proposes
ShotPlan targets cinematic video generation, where single-shot generation is not enough and coherent multi-shot structure matters.
- It adds learnable planning tokens that capture shot-level transition cues.
- These planning tokens are integrated with the original video generation tokens to control transition timestamps.
- The method uses Fractional Temporal Rotary Position Embedding (FRoPE) so transitions can be modeled at the frame level.
Reported outcome
The authors say ShotPlan outperforms existing cinematic video generation methods, with:
- more flexible shot management
- stronger inter-shot consistency
In short, the work is about making video generation behave more like a directed sequence of shots rather than a single uninterrupted clip.
More from Multimodal
- TimeLens2 claims SOTA on 7 video grounding benchmarks with 4B and 8B models — _akhaliq · 2026-07-21
- AI-made 4-minute horror short ‘THE NOT KNOW’ lands as a shareable demo — gen_ericai · 2026-07-21
- SVG Generation Comparison: Leading AI Models Draw a Red Ferrari — Able-Line2683 · 2026-07-21
- Adding order metadata makes VLM error detection collapse, new benchmark shows — m_wulfmeier · 2026-07-21
- Claude AGI Agent starts paging itself in Slack with a no-heartbeat alarm — Sauers_ · 2026-07-21
- Gemini Omni is being called a video-editing leap on par with Nano Banana — CodeByPoonam · 2026-07-21