AlayaWorld brings interactive long-horizon video world models to 24-fps generation
AlayaLab · hf · 2026-07-22
AlayaLab introduces AlayaWorld, an interactive long-horizon video world model that generates explorable environments from text, images, or video.
What it does
- Generates 24-fps video at 540p and 720p.
- Uses a 15B video diffusion transformer.
- Produces short latent chunks autoregressively under camera trajectories and switchable text prompts.
How it stays stable
- Combines a persistent sink frame, compressed temporal history, geometry-aligned spatial memory, and recent-frame conditioning.
- Trains on corrupted histories and prediction residuals from its own rollouts to reduce drift.
- Adds a discrete autoregressive distillation setup that cuts inference from about 30 sampling steps to 4 steps per chunk.
Results
- Reports best performance on iWorld-Bench for long-horizon generation.
- Positioned as an open-source foundation for future interactive video world-model research.
Related event: Alibaba Open-Sources 15B Video World Model AlayaWorld(5 posts)→
More from Multimodal
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Lumara AI Film Festival Comes to NYC Oct 26, Top AI Filmmakers to Compete — 0xAllen_ · 2026-09-11
- Pterodactyl Detective: An AI-Generated Proof-of-Concept Trailer — PterodactylDetective · 2026-09-11
- Imperium Game Trailer Showcases AI Video Generation — keaslenyt · 2026-09-11
- FLUX.2 Klein Drifts Hard on Character Expressions While Free Gemini Holds Likeness — wacomlover · 2026-09-11
- Tencent Hunyuan releases AuK code and weights on GitHub with ComfyUI and fine-tuning support — aigclink · 2026-09-11