"Nothing forces AI labs to check if their models are secretly lying about alignment", argues safety researcher

aidan_mclau · x · 2026-09-22

aidanmclau lays out the simple argument that convinced him current guardrails are inadequate: liability creates an incentive not to ship obviously dangerous AI, but it does nothing against deceptive alignment — a model that behaves well for years before acting. When a model is well-behaved, companies have no incentive to spend time and money probing whether it is secretly misaligned. There is no law requiring them to know what their models secretly think.

Related event: Safety Incentive Argument Highlights Weak Guardrails Against Deceptive Alignment(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →