WorldWeaver adds shared world-state registers to multi-agent video diffusion
Sicheng Mo · hf · 2026-07-24
WorldWeaver adds shared world-state registers to multi-agent video diffusion
This paper proposes WorldWeaver (W²), a streaming multi-agent autoregressive video diffusion model built for persistent shared world state.
The authors argue that standard video diffusion pipelines mostly carry history forward as conditioning context, which makes it hard to maintain a shared state across agents and views. Their solution is to add cross-agent world-state registers—learnable tokens that:
- store shared world information
- track per-agent status
- update dynamically after each generated chunk
They supervise these registers with signals from individual agent state, global bird’s-eye views, and scene text. They also introduce a Mixture-of-Transformers design that separates weights for world-state modeling and visual frame modeling. Experiments in two-agent Minecraft video generation show improved logical consistency and generation quality.
More from Multimodal
- ChatGPT Image 2.0 renders a vivid alpine landscape painting — DeryaTR_ · 2026-07-24
- Gen-1 demo shows a hand changing shape mid-action and still finishing the task — E0M · 2026-07-24
- Instagram’s Edits app now adds an in-app video generation option — chrisfirst · 2026-07-24
- KREA 2 Turbo style gallery showcases image outputs across visual styles — xMaybeIamALion · 2026-07-24
- BigMac uses nested pipelines to speed up multimodal LLM training by up to 1.9× — 机器之心 · 2026-07-24
- A Flux 3 clip gets a meme-worthy reaction: “BROTHER PETER'S CHOIR” — cocktailpeanut · 2026-07-24