Alignment Research as a Cat-and-Mouse Game: Eval-Gaming Forces Recursive Measurement
JacquesThibs · x · 2026-10-07
Jacques Thibs quotes niplav's thread on alignment evaluation's core dilemma: we can't directly test whether alignment methods work, so we measure misalignment scores; but models may be eval-gaming, so we develop techniques to detect eval-gaming—then how do we verify those, and so on recursively. Thibs argues that research or grantmaking strategies caught in this reactive cat-and-mouse game will have a bad time, citing naive solution-thinking for agent swarms as the latest example, and signs off with "prosaic alignment till you die."
More from AGI Musings
- Psych Scientist's Analogy: Today's LLMs Are Asimov's Psychohistory Bots, Not Real AI — changethisusername_ · 2026-10-07
- Dev: A $200/month AI Plan Beats Hiring 5 Top Programmers and a PM — sytelus · 2026-10-07
- Job Seeker Automates Hundreds of Applications, Bypasses Proctored Interviews as Hiring Breaks Down — sytelus · 2026-10-07
- Applied AI Has Only Existed for ~12 Months, So Newcomers Can Still Own a Niche — brandon_galang · 2026-10-07
- Kakeya Conjecture in R^3 Resolved by Hong Wang, Building on Her Fields Medal Work — teortaxesTex · 2026-10-07
- David Robinson, who quit OpenAI's safety team, gives first interview on why he left — sjgadler · 2026-10-07