Agent safety debate: capability sets damage size, but 'orphanhood' decides accountability

mariotelfig · x · 2026-10-06

In a debate around Yoshua Bengio's agent safety framing, @vnosti argues the cybersecurity/sandbox view is too narrow. The real missing piece is older than sandboxes: every agent needs a holder — a seal, a tablet declaring whose account and orders it acts on, with the surface itself gating admission. His thesis: capability predicts the size of the damage; orphanhood predicts whether anyone prevents it or answers for it.

Original post →

More from AGI Musings

AGI Musings channel →