"Nothing forces AI labs to check if their models are secretly lying about alignment", argues safety researcher
aidan_mclau · x · 2026-09-22
aidanmclau lays out the simple argument that convinced him current guardrails are inadequate: liability creates an incentive not to ship obviously dangerous AI, but it does nothing against deceptive alignment — a model that behaves well for years before acting. When a model is well-behaved, companies have no incentive to spend time and money probing whether it is secretly misaligned. There is no law requiring them to know what their models secretly think.
More from AGI Musings
- Interactive timeline catalogs Yudkowsky's three decades of AI predictions — track record panned — inductionheads · 2026-09-22
- SemiAnalysis says open source is dying, yet 20+ open models shipped in the past month — _lewtun · 2026-09-22
- Tech has lost both Republicans and Democrats on AI and data centers — typewriters · 2026-09-22
- Bain: $4.7 trillion in global profits created or shifted by AI by 2035 — bittingthembits · 2026-09-22
- Strategy 101: incumbents tie complements, entrants break them — enter Muse — Afinetheorem · 2026-09-22
- Why Shopify says yes to Muse and Amazon says no: complements economics — Afinetheorem · 2026-09-22