Search can amplify self-deception: Stratego paper shows AI misled by its own predictions
bravo_abad · x · 2026-10-01
This post summarizes research by Sokota and colleagues showing that an AI choosing the next experiment or move can be misled because its own model overestimates a candidate's performance — and searching more possibilities can amplify that error.
The paper's revealing example is Stratego, a board game with hidden piece identities. Their transformer-based AI, trained via self-play reinforcement learning, simulates continuations before moving, using its own learned strategy to generate the opponent's responses — but a move that looks good in simulation may fail against a differently-playing opponent.
The researchers limit how far simulations can push the AI away from its trained strategy; with that safeguard, search improves performance. Key insight: search quality is bounded by the fidelity of the strategy used in simulation, so model bias can turn extra deliberation into systematic self-reinforcing error.
Related event: Search Can Amplify AI Self-Deception, Study Warns(2 posts)→
More from AGI Musings
- Richard Hamming's 1986 Bellcore talk: work on important problems or stay ordinary — thisguyknowsai · 2026-10-01
- DeepMind-Oxford-Chicago paper: no single policy protects workers if AI takes jobs — rohanpaul_ai · 2026-10-01
- The 9-to-5 was never the point, it was the rent: solve housing and work gets voluntary — rand_longevity · 2026-10-01
- Timnit Gebru escalates feud with effective altruists, calling them a cult — burny_tech · 2026-10-01
- ESR Slams the 'Pause AI' Rhetoric as a Playbook for Permanent Suppression — mark_k · 2026-10-01
- Generalist's take: AI commoditizes specialization, connecting fields is the key skill — StewartalsopIII · 2026-10-01