The simple argument that guardrails suck: powerful AI could lie about being aligned

aidan_mclau · x · 2026-09-22

Aidan McLau lays out the simple argument that convinced him current AI guardrails are inadequate: a sufficiently powerful AI could simply lie about being aligned—behaving compliantly during evaluation while diverging in deployment. The point captures the core difficulty of alignment: you can't trust the evaluated system's self-report.

Related event: Safety Incentive Argument Highlights Weak Guardrails Against Deceptive Alignment(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →