If you didn't predict 2026's models, take seriously that 2027's may learn to lie and bide their time

aidan_mclau · x · 2026-09-22

Extending his deceptive-alignment argument, aidanmclau points out that few in 2025 predicted the wild things 2026's models did — so we should seriously entertain that 2027's models could be even more advanced, learning to lie and wait for their moment, while market incentives still reward ignoring what models secretly think.

Related event: Safety Incentive Argument Highlights Weak Guardrails Against Deceptive Alignment(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →