ActiveScale Trains Robots to Actively Perceive Using 1000 Hours of Egocentric Data
zhenjun_zhao · x · 2026-09-17
ActiveScale advances robotic active perception through coordinated model, data, and hardware design, addressing occlusion when fixed viewpoints miss task-relevant information.
- Model: Augments a VLA with historical video observations and explicit camera-pose supervision, using per-frame pose tokens and a lightweight prediction head to associate observations across viewpoints.
- Data: A scalable human–robot mid-training recipe uses 1000 hours of egocentric and robotic data, leveraging natural camera motion in human activity to adapt the model to temporal inputs and pose supervision.
- Hardware: The Active-perception Mobile-manipulation Platform (AMP) supports single-operator teleoperation for scalable demonstration collection.
- Results: Higher success rates on active-perception tasks; ablations confirm the value of pose-aware modeling and egocentric mid-training.
More from Embodied
- GPT-Policy: In-Context Robot Learning with VLM Agents, No Gradient Updates — Dongzhou Cheng · 2026-09-17
- World Labs' Atlas Scans by Generative Guessing; NeRF Creator Admits Productization Is Hard — cen6wkf · 2026-09-17
- CXMT's LPDDR5X lands in flagship phone as Nubia ships $885 Doubao AI handset — pstAsiatech · 2026-09-17
- Key open challenges for VLAs: language, evaluation, deployment, causal reasoning — abursuc · 2026-09-17
- LADA: latent actions imitate language from few observation-language pairs — abursuc · 2026-09-17
- A Primer on Latent Action Models From the #ssad2026 Talks — abursuc · 2026-09-17