EditWorld adds precise editing and flexible referencing to video world models, scoring 80.0 on editing
NanyangTechnologicalUniversity · hf · 2026-09-29
NTU's EditWorld extends video world models from exploration to precise modification, streaming editing instructions and reference images during autoregressive generation. Key pieces: Gated Causal Attention for temporally varying editing conditions, a Sparse Context mechanism for bounded long-horizon context, joint autoregressive + bidirectional training with annealed self-resampling, and a data synthesis pipeline for editing supervision. It also introduces WBench-Editing, where EditWorld achieves 73.8 overall and 80.0 editing scores, substantially outperforming existing methods. Code is open-sourced.
More from Multimodal
- Midjourney v8.2 prompt: silver gelatin darkroom-style morning black-and-white portrait — tisch_eins · 2026-09-29
- One prompt, no cuts: Kling 4.0 video realism is getting hard to distinguish from real footage — FinanceYF5 · 2026-09-29
- One prompt, no cuts: Kling 4.0 video realism is getting hard to distinguish from real footage — FinanceYF5 · 2026-09-29
- A cyberpunk city built entirely in code: thousands of towers, live-synthesized soundtrack, zero audio files — techartist_ · 2026-09-29
- Atlases Are Already Inside: Recovering Population Templates by Making Diffusion Models Collapse — kwangmoo_yi · 2026-09-29
- EvolvingAvatar Uses Test-Time Training to Make 3D Talking Heads Adapt as Conversations Unfold — HFUT-AI · 2026-09-29