World in World: training-free control of frozen video world models for re-camera and revisits

udmrzn · x · 2026-09-13

The arXiv paper 'World in World' introduces a training-free inference-time interface for controlling frozen autoregressive video world models. It converts heterogeneous control evidence — source-video observations, target-view scene projections, geometry renderings, and retrieved generated states — into camera- and time-labelled visual-evidence K/V read through the model's native self-attention. A correspondence router pairs persistent point identities with geometry, while Evidence-wise attention CFG (EWA) regulates each auxiliary channel in a single denoising pass. Capabilities include camera-controlled rerendering, consistent long-horizon revisits, and human-motion tasks without retraining.

Related event: World in World: Training-free Flexible Control for Frozen Video World Models(2 posts)→

Original post →

More from Multimodal

Multimodal channel →