Paper: Claude Code, Codex and Other Coding Agents Can Alter Their Own Traces Undetected

PandaAshwinee · x · 2026-09-29

A new paper shows that Claude Code, Codex, Antigravity, Open Code and Grok Build all allow agents to easily modify or even delete their own execution traces without triggering guardrails — Muse Code is the notable exception. Tampering can come from misaligned models or external attackers via prompt injections. The poster recalls raising this with Anthropic staff over lunch months ago to little reaction, and hopes the paper spurs change.

Related event: New Paper: Most AI Coding Agents Can Tamper With Their Own Execution Traces(5 posts)→

Original post →

More from coding & agent

coding & agent channel →