Four years ago Google's Phenaki generated multi-minute videos from changing prompts
dumierhan · x · 2026-10-06
Former Google researcher Dumitru Erhan looks back at Phenaki, released four years ago — an early text-to-video model supporting prompts that change over time and videos up to several minutes long.
- Its demos included a 2:28 first-person motorcycle story built from sequential prompts
- It supported image+prompt video generation and camera control
- The swimming teddy bear and Mars astronaut clips were landmark moments of early text-to-video
A nostalgic marker of how far video generation has come since Sora and Veo.
More from Multimodal
- RTX 5080 vs 5060 Ti benchmarked on FLUX Krea 2: 2.1x faster, wildly asymmetric LoRA penalty — Altruistic_Print717 · 2026-10-06
- Hedra Multiplayer launches: shared canvas where each teammate works with their own agent — henloitsjoyce · 2026-10-06
- Reka AI's 19B omni-model Rho-1 handles text, images, video and robot control in one net — The Decoder · 2026-10-06
- AI-made 'Watson and the Shark' to show at the National Gallery of Art for four months — Merzmensch · 2026-10-06
- Dreamina 2.0 goes canvas-first with Extract Motion, offers a month of Basic for $1.5 — HeyAmit_ · 2026-10-06
- Google's KeyRec Achieves Best Long-Video VLM Results With Just 10% of Visual Token Budget — google · 2026-10-06