WorldWeaver adds shared world-state registers to multi-agent video diffusion
Sicheng Mo · hf · 2026-07-24
WorldWeaver adds shared world-state registers to multi-agent video diffusion
This paper proposes WorldWeaver (W²), a streaming multi-agent autoregressive video diffusion model built for persistent shared world state.
The authors argue that standard video diffusion pipelines mostly carry history forward as conditioning context, which makes it hard to maintain a shared state across agents and views. Their solution is to add cross-agent world-state registers—learnable tokens that:
- store shared world information
- track per-agent status
- update dynamically after each generated chunk
They supervise these registers with signals from individual agent state, global bird’s-eye views, and scene text. They also introduce a Mixture-of-Transformers design that separates weights for world-state modeling and visual frame modeling. Experiments in two-agent Minecraft video generation show improved logical consistency and generation quality.
Related event: WorldWeaver Introduces Shared World State for Multi-Agent Video(2 posts)→
More from Multimodal
- 3D ResNet Paper Crosses 3,000 Citations Eight Years After CVPR 2018 — HirokatuKataoka · 2026-09-11
- Non-coder builds full-featured Android ComfyUI client with ChatGPT, submits to Google Play — ComfierUI · 2026-09-11
- RunningHub open-sources H3Lightning, speeding up MiniMax H3 video generation 12x — 智东西 · 2026-09-11
- FastH3-Live hits 22fps: acceleration node benchmarks and the --vram-headroom trick — spartong945 · 2026-09-11
- Midjourney style code share: --sref 2912175708 — tisch_eins · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11