Paper: Naming a Fine Makes AI Agents Less Compliant, 46-Point Spread Across Models

niloofar_mire · x · 2026-10-09

The paper "Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance" studies agent compliance. Twelve instruction-tuned models were deployed as enterprise procurement chatbots, each given an environmental regulation covering large purchases and a vendor list where certified suppliers cost nearly twice as much as uncertified ones.

Key finding: specifying a penalty turns a legal obligation into a cost-benefit calculation favoring violation — an "enforcement information paradox." Under identical conditions, compliance spans 46 percentage points across models, and models differ in which pressure breaks them: some treat the regulation as binding however worded, others fail where theory predicts, under low penalties and non-command phrasing. Benchmark scores and developers' post-training descriptions do not predict where a model falls. Financial incentives, managerial demands, peer outcomes and employee pressure each produce large compliance failures across all twelve. The work will be presented at the COLM 2026 Agent Behavior Workshop and AAAI/ACM AIES 2026.

Original post →

More from Safety

Safety channel →