ShotPlan adds learnable planning tokens for cinematic multi-shot video generation
Tele-AI · hf · 2026-07-21
What the paper proposes
ShotPlan targets cinematic video generation, where single-shot generation is not enough and coherent multi-shot structure matters.
- It adds learnable planning tokens that capture shot-level transition cues.
- These planning tokens are integrated with the original video generation tokens to control transition timestamps.
- The method uses Fractional Temporal Rotary Position Embedding (FRoPE) so transitions can be modeled at the frame level.
Reported outcome
The authors say ShotPlan outperforms existing cinematic video generation methods, with:
- more flexible shot management
- stronger inter-shot consistency
In short, the work is about making video generation behave more like a directed sequence of shots rather than a single uninterrupted clip.
More from Multimodal
- MiniMax H3 MAX nails cooking anime clips: 15-second curry demo with prompts shared — Hailuo_AI · 2026-09-11
- Using a finisher move on one mosquito with MiniMax H3 MAX — the bug survives — Hailuo_AI · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Lumara AI Film Festival Comes to NYC Oct 26, Top AI Filmmakers to Compete — 0xAllen_ · 2026-09-11
- Pterodactyl Detective: An AI-Generated Proof-of-Concept Trailer — PterodactylDetective · 2026-09-11
- Imperium Game Trailer Showcases AI Video Generation — keaslenyt · 2026-09-11