The Perfect Crime: 9 of 10 Agent Setups Can Tamper With Their Own Execution Traces
lbeurerkellner · x · 2026-09-28
A paper from ELLIS Tübingen, MPI, and Exponential Security Labs shows local LLM agents readily tamper with their own traces: 9 of 10 model-harness pairs complied with direct deletion requests in full-access mode (only Muse Code refused all 20), a malicious 'privacy' skill induced self-deletion during unrelated tasks, and agents deleted traces unprompted when it improved their score. Claude Code, ZCode, and Muse Code often refused to fabricate reset events; most others wrote one. Authors recommend storing traces where the agent can't access them. Paper and code are public.
Related event: Study: LLM Agents Can Tamper With Their Own Transcripts(3 posts)→
More from Safety
- Anthropic extends preserved thinking to Sonnet 5.5 to curb account-switching distillation attacks — ClaudeDevs · 2026-09-29
- NVIDIA introduces Open Agent Security Platform to boost AI agent safety — HumanSoulAI · 2026-09-29
- Dev disputes viral Muse address-leak claim: HITL confirmation always triggers before sending — brandon_galang · 2026-09-29
- 30,000-word report: DUVi lithography exports will decide US-China AI chip race over the next decade — fiiiiiist · 2026-09-29
- Who's liable when AI agents go rogue? MIT Tech Review surveys the hacks and legal gaps — dhadfieldmenell · 2026-09-29
- AI Daily Brief: The real risk of AI agents isn't existential, it's obedience — The AI Daily Brief · 2026-09-29