LIFT: video generation with joint camera and future-layout control via on-policy self-distillation

_akhaliq · x · 2026-10-09

New paper LIFT introduces a unified video generation framework controlling both camera motion and semantic-spatial composition of newly revealed regions using only a last-frame layout. Uses dual-mode on-policy self-distillation with a dense spatiotemporal layout teacher, and curates LIFT-Vista, a dataset with large viewpoint changes and joint camera/layout annotations.

Original post →

More from Multimodal

Multimodal channel →