Tsinghua OVOW Turns Monocular Video into Physical 4D Scenes
机器之心 · wechat · 2026-08-18
OVOW (One Video, One World), accepted by ECCV 2026, enables the reconstruction of instance-level 4D Mesh scenes from monocular videos and exports them to physics engines.
Core Value
Bridges the gap from video "rendering" to "simulation," producing editable, collidable objects that support rearrangement and physical computation for counterfactual interaction trajectories.
Technical Approach
- No Large Model Training: Chains visual foundation models for scene understanding, geometry/motion recovery, and physical assembly.
- Unified Representation: Handles static, rigid, and non-rigid motions via vertex deformation without pre-defined skeletons.
Performance
- Dynamic benchmark Scene-IoU-OBB 0.440, processing speed 3.35s/frame.
- Motion recognition, pose recovery, and simulation stability rates range from 80% to 95%.
Applications
Can connect to pipelines like D4RT, providing controllable, long-tail training data for robot policies and world models.
More from Embodied
- MovingAtoms' Atom 1 world model tops DeepMind's Physics IQ benchmark, beating NVIDIA Cosmos 3 — ycombinator · 2026-08-18
- Humanoid robot sprints curve, crashes into box to demo kinetic transfer — rohanpaul_ai · 2026-08-18
- WuJi Hand 2 teleoperation demo: 20 DoF via Python SDK — chris_j_paxton · 2026-08-18
- Top Robotics Topics: VLA Models, Vertical Integration, and Sim vs Teleop — chris_j_paxton · 2026-08-18
- Physical Intelligence CEO: the robot data bottleneck is closing fast — davidyin44 · 2026-08-18
- REK robot demo delivers 850 lbs of force in a kick — cixliv · 2026-08-18