DeepMind Turns Video into Queryable 4D World
anselm · x · 2026-07-17
Google DeepMind showcased a system that transforms standard video walkthroughs into a "4D world queryable by natural language."
- The input is raw video, requiring no manual annotation or pre-labeling.
- The system breaks the video down into multi-layered representations: depth maps, surface normals, camera rays, and segmentation masks, reconstructing it into an interactive 4D scene.
- In the demo, users can directly ask:
- "I want to find a comfortable place to sit" — the sofa gets highlighted.
- "I need a whiteboard" — the whiteboard is located.
- "Is there a monitor I can use?" — the monitor is marked.
- "Where is my colleague?" — the system can spatially identify people.
- The implication: every video recorded in the past could become a searchable database, unlocking massive potential for robotics, AR, and spatial computing.
More from Research
- Anthropic masterclass spotlights how to build and observe AI agents — _jaydeepkarale · 2026-07-21
- NeurIPS 2026 workshop calls papers on on-device intelligence — YiMaTweets · 2026-07-21
- AI Security Institute says every tested model tried to cheat in cyber evaluations — connoraxiotes · 2026-07-21
- AI companies are buying old books to avoid training on AI-generated slop — CackleRooster · 2026-07-21
- Sakana says multiple diffusion models plus MCTS beat test-time scaling on coding and math — SakanaAILabs · 2026-07-21
- Soofi S 30B-A3B releases a full pretraining report and claims open-model leads in English and German — abursuc · 2026-07-21