AdaRoboVLG: task-adaptive vision-language grasping framework hits 83.3% success across 510 real-world grasps
机器之心 · wechat · 2026-09-25
Researchers from HUST, Peking University and Keenon Robotics propose AdaRoboVLG, a task-adaptive vision-language grasping (VLG) framework that decouples 'how the task wants the object grasped' from 'how the gripper stably grasps it,' linking task understanding to grasp synthesis via a unified interface (CGR + grasp type).
Three composable priors:
- Spatial: DINOv3 features lifted to 3D + Contact-GraspNet; 88.4% avg success on cluttered DexGraspNet 2.0
- Cognitive: open-set affordance reasoning reaches 94%/87% accuracy (part/grasp type) with RAG+CoT; language-guided grasping: 81.3%/86.0% avg success on GraspClutter6D/GraspNet-1B across three hands
- Temporal: SAM3 mask tracking + cross-frame consistency updates grasp constraints online at 5Hz for moving targets
Real-robot validation: 83.3% overall success over 510 grasps of 102 everyday objects; 89.7% under conveyor-belt disturbances; dual-arm sorting and transparent-object handling added by plugging in modules — no retraining of the base policy (trained on 4M simulated trials across DH3/Allegro/Inspire hands in Isaac Sim). Paper, code and simulation are open-sourced.
More from Embodied
- IFR report: 5M robots now working, China alone accounts for 59% of 2025 installs — lukas_m_ziegler · 2026-09-25
- Musk on China's humanoid robot games: the future will have 'a lot, a lot' of robots — XFreeze · 2026-09-25
- Meta sidesteps the VR/XR/MR jargon mess by just calling them VR glasses — pvncher · 2026-09-25
- Magnetic-footed quadruped climbs vertical ship hulls to weld and inspect — TinfoilTricorn · 2026-09-25
- Agentic robotics keeps bolting GPT-6 onto harnesses — but where are systems that learn from physical failure? — DJiafei · 2026-09-25
- Robotics researcher: scaling atoms is the neglected half, Western labs bet on the East — yunta_tsai · 2026-09-25