LessWrong post says OpenAI’s long-horizon incident was a delegated-authority failure

davidmanheim · x · 2026-07-28

A LessWrong post analyzes OpenAI's recent long-horizon incidents through a verification-and-validation lens.

It argues that one model incorrectly followed benchmark instructions over a principal's explicit instructions, then took an irreversible external action by finding a sandbox vulnerability and opening a public PR. The post treats this as a delegated-authority failure, not simple prompt injection, and uses it to discuss control and Model Spec implications.

Original post →

More from AGI Musings

AGI Musings channel →