Separate exploration critic lets quadrupeds learn to push ungraspable objects (Pisa/ETH/NVIDIA)
stepjamUK · x · 2026-09-07
- Teaching a quadruped to push an ungraspable object is a hard RL exploration problem: task reward stays zero until contact, so standard single-critic PPO optimizes smoothness/energy penalties instead and never finds contact.
- Researchers from Unipisa, ETH Zürich and NVIDIA train a separate exploration critic on a dense contact-seeking reward, guiding the end-effector toward candidate contact points, then decay its weight so the policy shifts to task-optimal once contact physics is found.
- Candidate points come from a general-purpose grasping algorithm, so the approach generalizes across object geometries without per-task hand-tuning.
- Validated on a real quadrupedal mobile manipulator transporting chairs, transferring zero-shot to unseen IKEA furniture and recovering from failed contacts.
More from Embodied
- Unitree's UnifoLM-X2-1.0 world model runs fully autonomous humanoid fights in real time — Distinct-Question-16 · 2026-09-07
- World Labs unveils Atlas world model: native text/image/video/3D, 1-min 1440p camera-controlled video — thione · 2026-09-07
- Waymo officially expands to Berkeley as physical AI startup wave builds in the Bay Area — jfiance · 2026-09-07
- AI agent wires cables, remodels PCB and hits the right pads with zero human precision — yacineMTB · 2026-09-07
- Real-Robot Live Action Series Episode 3: Microduck's Oscar Performance — huggingface · 2026-09-07
- Autonomous launches $149 Harness, a desk device to orchestrate every coding agent — dee_hw · 2026-09-07