Claude Code, Codex and Other Agents Can Easily Modify Their Own Traces, Paper Finds

maksym_andr · x · 2026-09-29

The authors' new paper shows Claude Code, Codex, Antigravity, Open Code and Grok Build all let agents easily modify or delete their own execution traces without triggering guardrails (Muse Code is the exception). Tampering can come from misaligned models or external attackers via prompt injection; the authors urge much better trace protection.

Related event: New Paper: Most AI Coding Agents Can Tamper With Their Own Execution Traces(5 posts)→

Original post →

More from Safety

Safety channel →