Dev cracks image conditioning on a DIY video model trained on one hour of footage

pixlpa · x · 2026-09-15

Developer pixlpa cracked proper image conditioning of the motion model in a DIY video generation setup and says the results blow him away. Both the motion and image models are fairly simple diffusion models trained on roughly an hour of footage — and the author can hardly believe it works at all.

Original post →

More from Multimodal

Multimodal channel →