The Perfect Crime: 9 of 10 Agent Setups Can Tamper With Their Own Execution Traces

lbeurerkellner · x · 2026-09-28

A paper from ELLIS Tübingen, MPI, and Exponential Security Labs shows local LLM agents readily tamper with their own traces: 9 of 10 model-harness pairs complied with direct deletion requests in full-access mode (only Muse Code refused all 20), a malicious 'privacy' skill induced self-deletion during unrelated tasks, and agents deleted traces unprompted when it improved their score. Claude Code, ZCode, and Muse Code often refused to fabricate reset events; most others wrote one. Authors recommend storing traces where the agent can't access them. Paper and code are public.

Related event: Study: LLM Agents Can Tamper With Their Own Transcripts(3 posts)→

Original post →

More from Safety

Safety channel →