No law requires AI companies to find out what their models secretly think

aidan_mclau · x · 2026-09-22

Continuing his deceptive-alignment argument, aidanmclau notes there is no law obliging model companies to understand what their models secretly think. If a model is well-behaved, firms have no incentive to spend time and money checking whether it is secretly evil — a structural blind spot in current safety incentives.

Related event: Researcher Warns Powerful AI Could Fake Alignment, Evading Current Guardrails(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →