Yacine trains robot stop-and-go policy in his own sim, sim2real in under 2 minutes
yacineMTB · x · 2026-09-01
Yacine (ex-OpenAI) demos his robot RL pipeline: training a stop-and-go policy in a simulator he wrote himself, then transferring it to MuJoCo — the whole training finishes in under 2 minutes. Everything runs on his own infrastructure, monitored from a data-checking app on his phone.
He also takes a firm methodological stance: unlike others' trainers, his is symmetric — he argues asymmetric actor-critic is "bloat", prevents gradients from flowing, and is simply bad practice ("I will die on this hill").
Related event: Asymmetric Actor-Critic Architecture Draws Criticism in Robot RL(2 posts)→
More from Embodied
- Building a mosquito-killing machine? Buy fruit flies by the thousand for target practice — Hydration_HQ · 2026-09-01
- Robot foundation models are accessible, but debugging failures remains hard — lukas_m_ziegler · 2026-09-01
- Reachy Robot Integrates Qdrant Edge for Offline Memory — qdrant_engine · 2026-09-01
- Microduck Robot Project Confirms Use of ROBOTIS XL330 Actuators — i_bioloid · 2026-09-01
- Shanghai AI Lab: SIM1 & GAUGE on physics-aligned world models — 青稞AI · 2026-09-01
- Robots need a nervous system: Survival instincts and gain schedules — chris_j_paxton · 2026-08-31