Robotics RL is brutally hard: one practitioner's list of a dozen failure modes
Scobleizer · x · 2026-09-19
A robotics RL practitioner's candid retrospective: pattern matching isn't enough. Their run improved 2% then crashed to 100% failure — the KL term too low for the learning rate, the critic's gradient too high. The physics engine was too inaccurate for grasping, tactile sensors produced NaNs, sensors made training painfully slow, and sim frequency mismatched real-world inputs raising sim2real doubts. At one point entropy was zeroed by a simple reward bug, making "do nothing" optimal for three days before anyone noticed. Add an fp32 vs fp64 replay discrepancy, plus days lost to wrong answers from their AI assistant.
More from Embodied
- Fruit fly brain connectome drives an eBay Vector robot with 166,700 simulated neurons — sull · 2026-09-19
- Over 100M Americans Wear Sensors — But Do HRV and Readiness Scores Hold Up? — EricTopol · 2026-09-19
- Indie robot U-BOT day 32: prototype control board mount ready for motor torque testing — _Stocko_ · 2026-09-19
- Tesla AI5 chip enters trial production on Samsung's 2nm Texas fab, mass output by 2027 — XFreeze · 2026-09-19
- Real world isn't a simulator: why autonomous AI struggles to go physical — AlexTensor · 2026-09-19
- RoboHarm benchmark finds GPT-6 Astra and Claude Fable rarely refuse dangerous robot commands — The Decoder · 2026-09-19