Superhuman Stratego for a Few Thousand Dollars: Ataraxos Beats the Best Human Ever 15-4-1

Superhuman AI for Stratego Using Self-Play Reinforcement Learning and Test-Time Search

Samuel Sokota, Eugene Vinitsky, Hengyuan Hu, J. Zico Kolter, Gabriele Farina

cs.LG, cs.AI

2025-11-11

Self-play RL plus test-time search, trained for a few thousand dollars, beats the most decorated Stratego player in history 15-4-1 over 20 games (85% effective win rate).

What problem this solves

Stratego is a 10×10 board wargame where each player secretly arranges 40 pieces; piece identities stay hidden until pieces clash in battle. Among the classic games treated as major AI benchmarks, it has by far the most hidden information. Texas hold'em has 1,326 possible starting hands, few enough to enumerate; Stratego has more than 10^33 piece configurations. The public-information transformations behind Libratus and DeepStack are the most successful approach to imperfect-information games, but their cost scales with the amount of hidden information, and at this scale they cannot be applied. The 2022 Science effort on Stratego spent compute with a commercial cost in the millions of dollars and still fell short of top humans, which left Stratego as perhaps the only classic game where a well-resourced attempt failed to produce superhuman play.

Method

Ataraxos trains tabula rasa from self-play alone, with no human game data. Three decisions carry the result.

Training data comes from sampling the policy networks directly; AlphaZero-style search-based generation was tried and the slowdown outweighed the benefit, possibly because covering an enormous distribution of setups rewards game quantity over quality. Move samples are filtered to the top quartile by advantage magnitude, which cut wall-clock time per iteration by about 2.5x while increasing sample efficiency and final strength. Setups train on Monte Carlo returns, unusual in RL and also observed in LLM reasoning training. On the engineering side, a custom CUDA C++ simulator sustains about 10 million state updates per second on one H100, the replay buffer lives entirely in GPU memory, and bfloat16 adds another 3x.

Results

SettingMetricResult
vs Pim Niemeijer, 20 gamesRecord15 wins, 4 draws, 1 loss (85% effective win rate)
2025 World Championship demo, 40 gamesRecord38 wins, 2 losses, 0 draws (95%)
Policy network aloneElo2095
Depth-40, 1000-rollout searchElo2218, 1.26 s per move

Pim Niemeijer holds 4 world championships, 15 Dutch national titles, 2 online world championships, and over 600 weeks at world #1; a fellow player calls him the best Stratego player ever. The effective win rate counts draws as half wins, and 85% is a margin without precedent at the top level of human play. The format was also asymmetric: only Pim could adapt across games, which three-time world champion Vincent de Boer called a large handicap for the AI. Under an i.i.d. assumption the p-value is below 0.00026; a complementary betting analysis has an Ataraxos-backing bettor finish with more than 2,000x the wealth of the best comparator. Training used 16 H100s for one week for the policy networks and 4 H100s for four days for the belief network, priced by the authors at a few thousand dollars, against millions for the prior effort. In total: 163 million completed games, 208 billion environment steps. Removing the reverse KL term toward the move network during search drops Elo to 1733, below the no-search policy: the search overfits quirks of the network that other opponents do not reward.

Why it matters

The headline is cost. For strategic decision problems where a fast simulator can be built, the compute threshold for superhuman play fell from an industrial budget to an academic one. The components are generic: dynamic damping depends on nothing specific to Stratego, and the search is a single regularized policy update rather than a bespoke algorithm. Two pieces transfer directly to other RL projects: filtering training samples by advantage magnitude, and implementing decision-time search as a damped update. The honest caveat: this is a single-game demonstration, and the claim that many strategic settings are now cheaply solvable is the authors' conjecture, not something the experiments test.

Limitations

Stated by the authors: Elo is “useful but flawed” for Stratego because randomization compresses margins; the 20-game p-value leans on an i.i.d. assumption that human adaptation breaks; the failure of search-based data generation has only a speculative explanation; the benefit of advantage filtering merits further investigation; all ablations are single-seed. Beyond that: the top-tier human sample is one player; 60 evaluation games is small; the two losses in the demo show the system can be beaten; the belief network's robustness to humans playing far off the self-play distribution rests on dropout; and “a few thousand dollars” is the authors' accounting, with compute listed but no itemized bill.

Terms

Source

What people are saying

Related papers

All paper explainers