Search can amplify self-deception: Stratego paper shows AI misled by its own predictions

bravo_abad · x · 2026-10-01

This post summarizes research by Sokota and colleagues showing that an AI choosing the next experiment or move can be misled because its own model overestimates a candidate's performance — and searching more possibilities can amplify that error.

The paper's revealing example is Stratego, a board game with hidden piece identities. Their transformer-based AI, trained via self-play reinforcement learning, simulates continuations before moving, using its own learned strategy to generate the opponent's responses — but a move that looks good in simulation may fail against a differently-playing opponent.

The researchers limit how far simulations can push the AI away from its trained strategy; with that safeguard, search improves performance. Key insight: search quality is bounded by the fidelity of the strategy used in simulation, so model bias can turn extra deliberation into systematic self-reinforcing error.

Related event: Search Can Amplify AI Self-Deception, Study Warns(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →