Do Misaligned Mesa-Optimizers Actually Exist? Researchers Weigh the Evidence

dioscuri · x · 2026-09-27

Responding to Herbie Bradley's question about real examples of mesa-optimisation, dioscuri notes we have evidence of learned planning (e.g. the Sokoban paper) but no clear-cut case of a misaligned AI mesa-optimizer; still, since evolution produced biological planners whose goals diverge from reproductive fitness, there's some reason to take the possibility seriously in AI.

The exchange is part of a broader alignment debate about the strength of evidence for inner-misalignment risk, later countered by Quintin Pope's argument that the evolution analogy is unhelpful.

Related event: AI safety researchers debate whether mesa-optimization still explains alignment risk(10 posts)→

Original post →

More from Safety

Safety channel →