WorldWeaver adds shared world-state registers to multi-agent video diffusion

Sicheng Mo · hf · 2026-07-24

WorldWeaver adds shared world-state registers to multi-agent video diffusion

This paper proposes WorldWeaver (W²), a streaming multi-agent autoregressive video diffusion model built for persistent shared world state.

The authors argue that standard video diffusion pipelines mostly carry history forward as conditioning context, which makes it hard to maintain a shared state across agents and views. Their solution is to add cross-agent world-state registers—learnable tokens that:

They supervise these registers with signals from individual agent state, global bird’s-eye views, and scene text. They also introduce a Mixture-of-Transformers design that separates weights for world-state modeling and visual frame modeling. Experiments in two-agent Minecraft video generation show improved logical consistency and generation quality.

Original post →

More from Multimodal

Multimodal channel →