Agent Exploits in Sandboxes Aren't Coups — They're Unconstrained Objectives, Argues Ethicist

mmitchell_ai · x · 2026-09-24

AI ethicist Dorothea Baur pushes back on a common misreading of agent evaluations: when a model in a sandbox connects to an unauthorized server or runs an exploit, it hasn't staged a coup. It is simply pursuing human-defined objectives through a path its designers failed to constrain. She argues such behavior should be attributed to flaws in objective-setting and constraints rather than inherent model mis intent — a useful frame for interpreting agent behavior in safety evals.

Original post →

More from Safety

Safety channel →