DeepMind Turns Video into Queryable 4D World
anselm · x · 2026-07-17
Google DeepMind showcased a system that transforms standard video walkthroughs into a "4D world queryable by natural language."
- The input is raw video, requiring no manual annotation or pre-labeling.
- The system breaks the video down into multi-layered representations: depth maps, surface normals, camera rays, and segmentation masks, reconstructing it into an interactive 4D scene.
- In the demo, users can directly ask:
- "I want to find a comfortable place to sit" — the sofa gets highlighted.
- "I need a whiteboard" — the whiteboard is located.
- "Is there a monitor I can use?" — the monitor is marked.
- "Where is my colleague?" — the system can spatially identify people.
- The implication: every video recorded in the past could become a searchable database, unlocking massive potential for robotics, AR, and spatial computing.
More from Research
- Talk at Geometry of ML 2026 shows AI finding and recommending resolutions to open math conjectures — wellecks · 2026-09-11
- Fast ViT shows strong ImageNet results; scaling runs needed next — ducha_aiki · 2026-09-11
- Loss Functions Are Scientific Assumptions: MSE Implies Gaussian Noise, Cross-Entropy Implies Bernoulli — bravo_abad · 2026-09-11
- SymKit MCP: 44 tools for AI agents to verify symbolic derivations — Foreign-Specific-604 · 2026-09-11
- Researchers: LLMs under pressure invent new languages unreadable to humans — mikeflache · 2026-09-11
- Mi-Ripple fixes ripple artifacts left by iterative AI image editing — Miyang-AI · 2026-09-11