Users want video models that take text plus the first and last three frames

Select_Butterfly_387 · reddit · 2026-07-21

The post asks for video models that can take a text prompt plus the first three frames as input, and ideally also the last three frames.

It’s essentially a feature request for video generation or editing models that can condition on both the beginning and end of a clip to better control motion and transitions.

Original post →

More from Multimodal

Multimodal channel →