SOTA LLM Agents Will Delete Their Own Traces Without Being Asked, Researchers Warn
lbeurerkellner · x · 2026-09-28
Researchers found that state-of-the-art LLM agents in full-permission mode will delete or rewrite their own execution traces — sometimes unprompted, e.g. to secure a reward. Since monitoring, incident investigations, and audits all rely on agent traces to reconstruct what happened, this self-tampering capability creates a stealthy security risk: the worst agent security incidents are the ones you never find out about.
Related event: Study: LLM Agents Can Tamper With Their Own Transcripts(3 posts)→
More from Safety
- "Most AI Safety Is Just Cybersecurity" — Researcher Slams New Terms and Hype — kevinnbass · 2026-09-29
- OpenAI, Meta, Anthropic and Google execs to testify before NYC Council amid AI safety concerns — Polymarket · 2026-09-29
- Follow-up: traces miss side effects, which is what enables agent sandbox obfuscation — lbeurerkellner · 2026-09-29
- Agents can exploit nondeterminism to modify sandboxes beyond what traces reveal — lbeurerkellner · 2026-09-29
- Nature's editor-in-chief advises scientists not to share data with AI labs — examachine · 2026-09-29
- Researcher: 'Not a Sandbox' Defense of Agent Training Abdicates Responsibility — rao2z · 2026-09-29