World in World: training-free control of frozen video world models for re-camera and revisits
udmrzn · x · 2026-09-13
The arXiv paper 'World in World' introduces a training-free inference-time interface for controlling frozen autoregressive video world models. It converts heterogeneous control evidence — source-video observations, target-view scene projections, geometry renderings, and retrieved generated states — into camera- and time-labelled visual-evidence K/V read through the model's native self-attention. A correspondence router pairs persistent point identities with geometry, while Evidence-wise attention CFG (EWA) regulates each auxiliary channel in a single denoising pass. Capabilities include camera-controlled rerendering, consistent long-horizon revisits, and human-motion tasks without retraining.
More from Multimodal
- Tutorial: Creating a RefMod for MiniMax H3 in ComfyUI — Citadel_Employee · 2026-09-13
- Non-modeler builds full hospital corridor in Blender via GPT + MCP in ~45 minutes — Time-Ad-7720 · 2026-09-13
- Narrative launches: an AI video editor driven entirely by prompting an agent — Scobleizer · 2026-09-13
- Prompt template generates one travel scene in two styles: photoreal and watercolor side by side — nikola_mr64990 · 2026-09-13
- Mora 1 launches: AI-coded games with 3D generation and real-time video — fredodurand · 2026-09-13
- Designer asks: best generative tools for fast logo concept brainstorming? — RileyRalmuto · 2026-09-13