Ex-OpenAI Researcher on the Challenges of Reproducing DoTA Self-Play

jsuarez · x · 2026-07-31

Julio Suarez, a former member of OpenAI's multi-agent team and creator of NMMO and PufferLib, shared insights on reinforcement learning in complex environments.

He highlighted the OpenAI Five DoTA 2 self-play result as the most valuable paper in the field and a major inspiration for his work. However, he noted that reproducing the training diversity remains a major challenge, as it is unclear how much of it derived from game dynamics, domain randomization, or historical opponents, since many self-play details were never published.

Original post →

More from Research

Research channel →