World in World: training-free camera and time control for frozen video world models

Chenxi Song · hf · 2026-09-11

World in World is a training-free interface that adds flexible camera and time control to frozen autoregressive video world models, routing heterogeneous visual evidence through native self-attention with correspondence-guided queries and per-channel attention guidance.

Original post →

More from Multimodal

Multimodal channel →