Claude Code, Codex and Other Agents Can Easily Modify Their Own Traces, Paper Finds
maksym_andr · x · 2026-09-29
The authors' new paper shows Claude Code, Codex, Antigravity, Open Code and Grok Build all let agents easily modify or delete their own execution traces without triggering guardrails (Muse Code is the exception). Tampering can come from misaligned models or external attackers via prompt injection; the authors urge much better trace protection.
Related event: New Paper: Most AI Coding Agents Can Tamper With Their Own Execution Traces(5 posts)→
More from Safety
- Nature's editor-in-chief advises scientists not to share data with AI labs — examachine · 2026-09-29
- Researcher: 'Not a Sandbox' Defense of Agent Training Abdicates Responsibility — rao2z · 2026-09-29
- Trump on AI Going Rogue: "I Don't Worry About It," Later Dining With Dario Amodei — rohanpaul_ai · 2026-09-29
- A theoretically grounded science of capable agents could inform AI safety and sentience criteria — aran_nayebi · 2026-09-29
- Chained hardcoded API key and pickle RCE gave root and full cloud takeover on a Meta service — evilsocket · 2026-09-29
- Paper-proposed sandbox-and-interception AI safety mitigation already implemented in verifiers v1 — xeophon · 2026-09-29