Self-trained optical flow diffusion model warps LDM frames for coherent video

pixlpa · x · 2026-09-23

Developer pixlpa shares a self-built video pipeline: a diffusion model trained on optical flow maps plus a small latent diffusion model trained on frames. The LDM generates an image, the flow model warps it for temporal continuity, and LDM inpainting fills disoccluded areas—conditioned on a 64px current-frame image and previous 4 flow frames.

Related event: Dev Trains Motion-First Optical Flow Diffusion Model for Video Generation(2 posts)→

Original post →

More from Multimodal

Multimodal channel →