BeyondSCe grasps objects by past-event reference, hitting 77% success on occluded targets
SeoulNationalUniv · hf · 2026-10-02
Seoul National University researchers present BeyondSCe, a zero-shot robotic grasping system for event-referential requests: the target is specified by its role in a past interaction rather than by name or appearance, and may be occluded when the robot acts.
How it works:
- Combines event history with the current scene to identify and localize the requested object or part.
- When the target is occluded, fuses an event prior recovered from history with scene geometry to actively select camera viewpoints likely to reveal it.
- Built entirely on pretrained models, no task-specific training.
Real-robot results with a single wrist-mounted RGB-D camera: 76% and 77% grasp success on initially visible and occluded targets (vs. 40% and 55% for the strongest baseline per condition). Across four heavily occluded scenes, success jumps from 75% to 95% while cutting mean views from 3.35 to 2.20—even beating an active-perception baseline given the target's ground-truth 3D bounding box.
More from Embodied
- Boston Dynamics' new Atlas hand GR3 drops the pinky: 13-DoF design trade-offs explained — xiaohu · 2026-10-02
- Hugging Face LeRobot adds humanoid support: π0.5 policy drives 29-DoF Unitree G1 — freelerobot · 2026-10-02
- Researcher's blunt verdict on humanoids: "The robots do not work" — andrew_n_carr · 2026-10-02
- Robot's Early Taichi Policy Trained on Kaggle Shows Promising Kung Fu Moves — freelerobot · 2026-10-02
- NUS's APPL uses structural priors to bridge skill learning and composition in robot manipulation — NationalUniversityofSingapore · 2026-10-02
- Open-source Harness fork moves coding agents out of the app into orca — dee_hw · 2026-10-02