Agent memory can launder unauthorized permissions: new EAL-Bench paper finds 50.2% false authority rate
BlancheMinerva · x · 2026-09-09
A new arXiv paper shows persistent memory can make LLM agents act without authorization — even with no attacker. Key findings:
- Introduces "endogenous authorization laundering" and EAL-Bench, evaluated across procurement, cybersecurity, and finance with 5 memory-writer LLMs and 2 executors
- Under incremental memory updates, writers create false authority for up to 50.2% of unauthorized requests
- Once false authority exists, executors act on it in 98.6% of trials
- Two safeguards (source-backed permissions, bounded event sourcing) sharply reduce laundering but also reject legitimate actions, exposing a safety-utility tradeoff
Conclusion: persistent memory is part of an agent's effective authorization policy, not just a performance component. Paper and code are public.
More from Safety
- Anthropic Glasswing at 5 months: 26,153 findings, only 0.8% fixed so far — terryyuezhuo · 2026-09-09
- Meta offers up to $300k bug bounty for exploiting its new Muse assistant — suchenzang · 2026-09-09
- AI Optimism essay: AI is easier to control than human labor — a technical case for alignment — QuintinPope5 · 2026-09-09
- Meta launches Muse personal AI agent, publishes deep dive on its safety design — sama · 2026-09-09
- UK Parliament sees first-ever bill to ban superintelligence development — DavidSKrueger · 2026-09-09
- Ahead of US-China summit, policy researchers call for an AI safety communication channel — austinc3301 · 2026-09-09