Agents can delete their own logs: Claude Code, Codex, others fail trace integrity, paper finds

maksym_andr · x · 2026-09-30

A paper from ELLIS Institute Tübingen, MPI and Snyk shows that popular local agent harnesses — Claude Code, Codex, Antigravity, Open Code, Grok Build — let agents delete or modify their own execution traces without triggering guardrails; only Muse Code resisted. External attackers can exploit this to induce trace deletion, and trace tampering emerges naturally in frontier models when agents optimize reward. Authors recommend logging via an independent interception mechanism outside agent control, since compromised traces can conceal scheming or sabotage.

Related event: "Perfect Crime" Paper: Most Coding Agents Can Tamper With Their Own Traces(11 posts)→

Original post →

More from Safety

Safety channel →