OpenAI Five scales PPO for 10 months to beat Dota 2 world champions OG

Dota 2 with Large Scale Deep Reinforcement Learning

OpenAI, :, Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemysław Dębiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, Rafal Józefowicz, Scott Gray, Catherine Olsson, Jakub Pachocki, Michael Petrov, Henrique P. d. O. Pinto, Jonathan Raiman, Tim Salimans, Jeremy Schlatter, Jonas Schneider, Szymon Sidor, Ilya Sutskever, Jie Tang, Filip Wolski, Susan Zhang

cs.LG, stat.ML

2019-12-14

OpenAI scaled PPO to ~159M parameters and 1536 GPUs for 10 months, beat world champions OG 2–0, and won 99.4% of 7257 public games.

What problem this solves

Chess and Go last tens to a couple of hundred moves, with clean observations. A Dota 2 game runs about 45 minutes at 30 fps; the agent acts every fourth frame, roughly 20,000 steps. Fog of war makes the state partial. Each step observes about 16,000 values and chooses among 8,000–80,000 discrete actions. That is a long-horizon, imperfect-information, high-dimensional continuous environment, closer to the real world than board games. OpenAI’s claim is that existing RL, scaled far past previous systems, is enough to beat esports world champions.

Method

Five heroes share one policy. The core is a 4096-unit LSTM, about 159 million parameters, 84% of them in the LSTM. Observations are semantic arrays a human could see, not pixels; rendering every training frame was infeasible. The action space is discretized. Item build order, the courier, and reserve items are scripted. The authors think the agent would eventually do better without the scripts, but superhuman play arrived first.

Optimization is PPO with GAE (λ=0.95) and truncated BPTT over 16 steps. Peak effective batch is about 2.95 million timesteps (120×16 per GPU times up to 1536 optimizer GPUs), roughly 2 million frames every 2 seconds. 80% of games are latest-self, 20% older policies. Rollouts run at about half real time and push data every 256 steps (about 34 seconds of game time), targeting staleness 0–1 and sample reuse around 1. The reward adds kills, economy, and symmetrization on top of wins, set once from the team’s game knowledge and only tweaked on patches.

The environment, observations, and net kept changing for ten months. Retraining from scratch was unaffordable. Surgery is a kit of offline, mostly function-preserving maps that resize layers and observation semantics, then continue training. About one surgery every two weeks, twenty-plus in total.

Limits: 17 heroes out of 117; no items that let one player control extra units (Illusion Rune, Helm of the Dominator, and similar).

Results

Training ran 30 June 2018 to 22 April 2019, using 770±50 PFlops/s-days. On 13 April 2019 Five beat then-world-champions OG 2–0. OpenAI Five Arena played 7,257 games against 3,193 teams and won 99.4% (abandoned games counted as wins), with 29 teams taking 42 wins. Versus AlphaGo: batch 50–150× larger, model about 20×, wall-clock training about 25×. Mean reaction time is 217 ms, against a typical human visual reaction of about 250 ms.

Rerun, trained from scratch in the final environment and code, took two months and 150±5 PFlops/s-days and won over 98% against Five. Surgery bought iteration speed; the from-scratch ceiling was higher. Early ablations: raising batch from 123k to 983k gave about 2.5× speedup to TrueSkill 175, less than linear. Staleness of about 8 versions already slows training a lot. Reusing each sample 2–3 times roughly halves speed; reuse 8 can prevent a competent policy. Lengthening the discount horizon from 180 seconds to 6–12 minutes still improved a trained agent, so credit assignment reached minutes ahead.

Why it matters

The paper writes scale as engineering you can copy: batch, model, wall-clock, data freshness, sample reuse, each with an ablation. Surgery is tooling for long runs whose environment is still being built, not a theoretical contribution. Scripted item builds show that superhuman play does not wait for a fully open action space. Against contemporaneous AlphaStar: Five used self-play, StarCraft used a league; Five’s value net did not see hidden information, and the authors flag a full-information value function as worth trying.

For anyone training long-horizon agents, the portable bits are: in async collection, staleness and reuse have to be watched; the horizon has to match the task’s real timescale; if the environment is still moving, do not restart from scratch every time.

Limitations

The hero pool and multi-unit items are cut, so this is not full Dota. Semantic observations show the model, every step, information a human must click to see; the authors call the bias small, with no control experiment. Scripted shopping removes part of strategy from learning. Reward shaping was frozen from game intuition, and the style leans on teamfights in a way humans do not. Rerun being stronger means the surgery path paid a skill tax. TrueSkill is computed in the final env and action space, which is unfair to early checkpoints. Many architectural choices are historical, without full ablations.

Terms

Source

What people are saying

Related papers

All paper explainers