World in World: training-free camera and time control for frozen video world models
Chenxi Song · hf · 2026-09-11
World in World is a training-free interface that adds flexible camera and time control to frozen autoregressive video world models, routing heterogeneous visual evidence through native self-attention with correspondence-guided queries and per-channel attention guidance.
More from Multimodal
- Turn any article into a podcast with Meta's Muse — alexandr_wang · 2026-09-11
- GPT-6 Astra 3D Workflow: Blender MCP for Hard-Surface, TripoAI for Organic Models — majidmanzarpour · 2026-09-11
- Before Diffusion: Looking Back at VQGAN, the Pre-Diffusion Foundation of Image Generation — makeitrad1 · 2026-09-11
- ComfyUI node brings a 3-light 3D dome relighting studio to MiniMax H3 Edit — Emotional_Example_12 · 2026-09-11
- Optimized video workflow: 10s at 1MP in ~125s on an RTX 5090, with custom audio driving — Tokyo_Jab · 2026-09-11
- Seedance 2.5 generates ultra-realistic 30s Seoul lifestyle video with full prompt released — SimplyAnnisa · 2026-09-11