S2PD: serial computation in high-noise diffusion makes video models follow physics
elliottszwu · x · 2026-10-07
Researchers from Cambridge and Toyota Motor Europe introduce S2PD (Serial-to-Parallel Diffusion), tackling why bidirectional video diffusion models violate physical laws and simple symbolic rules even when trained on massive in-distribution data.
- Key idea: run autoregressive (serial) diffusion at high noise to provide the serial computation needed to coordinate interdependent events and produce valid state transitions, then switch to parallel diffusion at low noise to jointly refine the video with faster sampling than fully serial methods.
- Two implementations: a pixel-space diffusion transformer trained from scratch, and a pretrained video model adapted via LoRA fine-tuning with causal attention.
- Results: across games, physical simulations, and real video, S2PD follows rules more reliably than matched bidirectional baselines, with better temporal stability and sampling efficiency than other serial approaches; a single start image yields multiple plausible continuations.
More from Multimodal
- Image Edit Arena Launches Multi-Image Edit Leaderboard: gpt-image-2.5 Takes Top Two Spots — arena · 2026-10-07
- Generating a finished car-chase film from low-poly previz: a full video-gen prompt workflow — Ror_Fly · 2026-10-07
- "Smooth Reimu" Video Showcases Striking AI Character Animation — Orichalchem · 2026-10-07
- Solo dev trains 1B world model running realtime on RTX 5090 with prompt-switching — lucidml_lover · 2026-10-07
- Local NVFP4 Qwen 27B replicates Opus-style code-rendered video, no API needed — yzjJosh · 2026-10-07
- Gemini 3 Pro image bills output at 60x input: one user's $253 lesson — Atm1n9 · 2026-10-07