APC architecture reduces AgentDojo data exfiltration to 0% in new safety paper

GoodMarch3690 · reddit · 2026-08-23

A new paper addresses how AI agents can achieve unauthorized outcomes by combining individually permitted actions. The proposed Agentic Principal Chain (APC) tracks delegated authority and enforces checks against previous actions outside the model. In evaluations involving 3,154 instances (InjecAgent, AgentDojo, ASB), APC reduced AgentDojo data exfiltration from 75-100% to 0% across all domains and blocked all 544 InjecAgent data-stealing cases.

Original post →

More from Safety

Safety channel →