KAIST's World Observer keeps video world models tracking objects beyond the actor's view
kaist-ai · hf · 2026-10-02
KAIST AI introduces World Observer, addressing a core weakness of actor-centric video world models: objects leaving the agent's view lose their state and dynamics upon re-entry.
Key ideas:
- Decoupling observing from acting: the model jointly generates an agent-centric actor plus freely placeable panoramic observers that keep watching selected regions, so out-of-view objects continue to evolve visually and re-enter with updated states.
- Geometric grounding: actor and observers are warped from a shared panoramic source for explicit correspondence, while an Observer Sink of high-resolution references restores fine appearance at re-entry.
- Controllability: observers accept control signals to steer out-of-view evolution.
The team also releases world-space metrics and a benchmark spanning real and synthetic scenes. World Observer substantially improves out-of-view dynamics while staying competitive on fidelity, camera control, and 3D adherence.
More from Research
- AMap open-sources ABot-Recon: streaming 3D reconstruction from video with a 12-frame local context — rsasaki0109 · 2026-10-02
- Neuralink pretrains decoders on 50,000+ hours of neural data—dataset may be the real moat — CurieuxExplorer · 2026-10-02
- CyberGym cybersecurity benchmark effectively saturated on verified task subset — aryaman2020 · 2026-10-02
- SemEval-2027 Task 9 calls for teams on 4-language multimodal news framing analysis — preslav_nakov · 2026-10-02
- Alternating prompt and model upgrades lift science agent from 42% to 73% — rohanpaul_ai · 2026-10-02
- ISMIR paper teaches a transformer to play in 12 jazz piano legends' styles — umpedronosapato · 2026-10-02