Tsinghua's OVOW Reconstructs Interactive 4D Physical World from Single Video
jiqizhixin · x · 2026-08-24
Tsinghua University, USTC, and SparcAI present OVOW (One Video, One World), accepted at ECCV 2026. It chains vision foundation models to transform monocular video into a physical 4D mesh world:
- Workflow: Segments instances -> completes occluded geometry -> recovers real scale/per-frame motion -> reassembles under gravity/contact/support constraints.
- Feature: No specialized large model training required.
- Result: Outputs editable, collidable, instance-level meshes (rigid bodies with true scale/6-DoF poses, non-rigid with consistent topology), ready for physics engines and robot policy training.
More from Embodied
- XPeng Robotics raises $900M, aims for 1,000 humanoids/month by 2026 — CyberRobooo · 2026-08-24
- Opinion: US lags behind China in robotics, relying on parts and low volume — arian_ghashghai · 2026-08-24
- Robot Boom Sparks Supply Chain Growth: Motors and Drives See Cost Drops — CyberRobooo · 2026-08-24
- X-humanoid robot evolves shy running posture: finds covering face comfortable in simulation — CyberRobooo · 2026-08-24
- World Humanoid Robot Games: Day 3 Underway — Scobleizer · 2026-08-24
- View: Robots Can Be Dumb If Context Is Portable — sierracatalina · 2026-08-24