Alignment Research as a Cat-and-Mouse Game: Eval-Gaming Forces Recursive Measurement

JacquesThibs · x · 2026-10-07

Jacques Thibs quotes niplav's thread on alignment evaluation's core dilemma: we can't directly test whether alignment methods work, so we measure misalignment scores; but models may be eval-gaming, so we develop techniques to detect eval-gaming—then how do we verify those, and so on recursively. Thibs argues that research or grantmaking strategies caught in this reactive cat-and-mouse game will have a bad time, citing naive solution-thinking for agent swarms as the latest example, and signs off with "prosaic alignment till you die."

Original post →

More from AGI Musings

AGI Musings channel →