Apple's STARFlow-V: First Normalizing Flow-Based Causal Video Generator Matches Diffusion
thoma_gu · x · 2026-09-01
Apple's team released STARFlow-V, the first normalizing flow-based causal video generator, showing flows can match video diffusion models in visual quality.
Key ideas:
- Operates in a spatiotemporal latent space with a global-local architecture that restricts causal dependencies to a global latent while preserving rich within-frame interactions, easing the error accumulation common in autoregressive diffusion generation.
- Proposes flow-score matching, adding a lightweight causal denoiser to improve autoregressive consistency.
- A video-aware Jacobi iteration scheme parallelizes inner updates without breaking causality, improving sampling efficiency.
- Thanks to its invertible structure, one model natively supports T2V/I2V/V2V, end-to-end training, and exact likelihood estimation.
Context: amid debates that most "autoregressive video" models are distilled bidirectional DiTs, this offers an alternative path. Paper and code are public.
More from Multimodal
- VideoDeltaNet open-sources hybrid attention that speeds up MiniMax H3 video generation up to 90x — realmrfakename · 2026-09-03
- jjk-explain turns any concept into a Jujutsu Kaisen-style explainer video with one Claude Code command — teortaxesTex · 2026-09-03
- New ComfyUI node adds WYSIWYG video cropping with 8 fixed ratios on the Load Video preview — MayaProphecy · 2026-09-03
- WAN 3 Tops AI Video Editing Chart, Ranked #1 With Audio at 1189 Elo — koltregaskes · 2026-09-03
- MiniMax H3 open-weight model generates character-drawing timelapses with cursor UI — WolframRvnwlf · 2026-09-03
- 'Reincarnated as a Vape': AI-generated short film leans into absurd premises — superfatbeagle · 2026-09-03