ReViV reconstructs egocentric 4D viewer-and-scene dynamics from one monocular video
ethz-vlg · hf · 2026-07-21
ReViV reconstructs both the viewer and the scene in 4D from one egocentric video
ETH Zurich researchers present ReViV, a unified framework for holistic egocentric 4D reconstruction from a single monocular RGB video.
- It jointly models RGB video, camera trajectory, gaze, full-body motion, hand motion, and depth as a full multimodal distribution.
- The system uses a Masked Generative Egocentric Transformer in a single feed-forward architecture, aiming to avoid separate pipelines for scene perception and ego-motion.
- Across HoloAssist, HOT3D, ARCTIC, Aria Digital Twin, and TACO, the paper reports state-of-the-art accuracy and efficiency for ego-body, hand, gaze reconstruction and camera tracking, while staying competitive on depth estimation.
- The code and models are open sourced at reviv4d.github.io.
More from Embodied
- Tesla expands Robotaxi rides to seven areas, including new Orlando and Tampa zones — elonmusk · 2026-07-22
- Hands-on robotics workshop on Saturday may be the last in-person session before August — StewartalsopIII · 2026-07-22
- NVIDIA pitches World Foundation Models as a way to scale physical AI data generation — MonaJalal_ · 2026-07-22
- RoboMME Podcast Preview: Benchmarking Memory for Robotic Policies — chris_j_paxton · 2026-07-21
- Gritt says an 8-person crew now installs 3,000 to 4,000 solar panels a day — HaktanSuren · 2026-07-21
- A helium-powered flying robot whale aims to be a quiet companion pet — chris_j_paxton · 2026-07-21