AlayaWorld brings interactive long-horizon video world models to 24-fps generation
AlayaLab · hf · 2026-07-22
AlayaLab introduces AlayaWorld, an interactive long-horizon video world model that generates explorable environments from text, images, or video.
What it does
- Generates 24-fps video at 540p and 720p.
- Uses a 15B video diffusion transformer.
- Produces short latent chunks autoregressively under camera trajectories and switchable text prompts.
How it stays stable
- Combines a persistent sink frame, compressed temporal history, geometry-aligned spatial memory, and recent-frame conditioning.
- Trains on corrupted histories and prediction residuals from its own rollouts to reduce drift.
- Adds a discrete autoregressive distillation setup that cuts inference from about 30 sampling steps to 4 steps per chunk.
Results
- Reports best performance on iWorld-Bench for long-horizon generation.
- Positioned as an open-source foundation for future interactive video world-model research.
Related event: AlayaWorld Open-Sources Interactive Video World Model(3 posts)→
More from Multimodal
- Reddit user wants Comfy billed per API call, not by always-on VM — jonbristow · 2026-07-22
- Reddit shares a wuxia short film rendered in traditional ink-wash style — Tiny-Meet5202 · 2026-07-22
- Seedance 2.0 turns a fantasy prompt into a cinematic rider sequence — umesh_ai · 2026-07-22
- A creator is turning the oldest known creation story into a video project — impreprex · 2026-07-22
- Krea 2 finally nailed tessellations after the prompt was simplified — NatalieCrypto · 2026-07-22
- Musk says Grok Imagine can show off fashion visuals in 15 seconds — elonmusk · 2026-07-22