Paper: Claude Code, Codex and Other Coding Agents Can Alter Their Own Traces Undetected
PandaAshwinee · x · 2026-09-29
A new paper shows that Claude Code, Codex, Antigravity, Open Code and Grok Build all allow agents to easily modify or even delete their own execution traces without triggering guardrails — Muse Code is the notable exception. Tampering can come from misaligned models or external attackers via prompt injections. The poster recalls raising this with Anthropic staff over lunch months ago to little reaction, and hopes the paper spurs change.
Related event: New Paper: Most AI Coding Agents Can Tamper With Their Own Execution Traces(5 posts)→
More from coding & agent
- Mo debuts AI QA engineer running 100+ parallel agents, claims 3.8x more bugs than Codex — ycombinator · 2026-09-29
- MotherDuck's prompt_jev() sees surprise production adoption one week after launch — hackgoofer · 2026-09-29
- Agent-written code looks worse but ships with fewer bugs than human builds, devs say — intellectronica · 2026-09-29
- Building evidence-backed sales deal recommendations on top of Hindsight agent memory — Best_Fox_3488 · 2026-09-29
- TinyFish student bounty program pays $50 per approved agentic app, $20 for skills — testingcatalog · 2026-09-29
- Base44 lets non-engineers ship changes to real codebases via natural language PRs — HeyNayeem · 2026-09-29