Cambridge's S2PD uses serial diffusion to make video generation physically consistent

kastnerkyle · x · 2026-10-08

Researchers from Cambridge and Toyota Motor Europe introduce S2PD (Serial-to-Parallel Diffusion), addressing why bidirectional video diffusion models violate physics and simple symbolic rules even when trained on unlimited procedural data.

Method: autoregressive diffusion at high noise provides the serial computation needed for valid state transitions and interdependent events; parallel diffusion at low noise jointly refines the video, cutting sampling time versus fully serial generation. Implemented both as a pixel-space diffusion transformer trained from scratch and via LoRA + causal attention on a pretrained video model.

Results: across games, simulations, and real video, serial methods beat matched bidirectional baselines, produce multiple plausible continuations from one start frame, and offer better temporal stability and efficiency than other serial approaches.

Related event: Cambridge's S2PD Makes Video Diffusion Physically Consistent(3 posts)→

Original post →

More from Multimodal

Multimodal channel →