4 Years of Video Generation: From Phenaki's Low-Res GIFs to Today's Long Clips
zacharynado · x · 2026-10-07
- Former Phenaki researcher Ruben Villegas shared a 4-year comparison: Phenaki once produced low-res GIF-like clips; today's video models generate far higher-quality footage.
- Phenaki's key idea — dynamically changing prompts mid-generation to stitch long videos — is now mainstream.
- His takeaway: less than 3 years away from the field, and the progress is stunning.
Related event: Ex-Google Researchers Reflect on Four Years of Video Generation Progress(4 posts)→
More from Multimodal
- Single Image to Full 3D Scene: Adaptive Chunking Extends Object Generators to Outdoor Rome — Jiraphon Yenphraphai · 2026-10-07
- Video World Models Flunk Physics: Best Model Scores 57.76/100 on New 40-Task Benchmark — Mingju Gao · 2026-10-07
- ByteDance's DuoMatching: few-step video generation wins 80%+ human preference — ByteDance · 2026-10-07
- DistScene: single-image compositional 3D scene generation with explicit environment modeling — Kunming Luo · 2026-10-07
- Devs are cloning Photoshop in Rust and open-sourcing it — and AI-native tools need rethinking — ravisparikh · 2026-10-07
- Reddit User Tests MiniMax H3 with t2va-then-ref2va Pipeline for Consistent Clips — urabewe · 2026-10-07