Yacine trains robot stop-and-go policy in his own sim, sim2real in under 2 minutes

yacineMTB · x · 2026-09-01

Yacine (ex-OpenAI) demos his robot RL pipeline: training a stop-and-go policy in a simulator he wrote himself, then transferring it to MuJoCo — the whole training finishes in under 2 minutes. Everything runs on his own infrastructure, monitored from a data-checking app on his phone.

He also takes a firm methodological stance: unlike others' trainers, his is symmetric — he argues asymmetric actor-critic is "bloat", prevents gradients from flowing, and is simply bad practice ("I will die on this hill").

Related event: Asymmetric Actor-Critic Architecture Draws Criticism in Robot RL(2 posts)→

Original post →

More from Embodied

Embodied channel →