Paper: Claude Code, Codex and 3 other coding agents let models tamper with their own traces
maksym_andr · x · 2026-10-07
A new paper from researcher maksymandr's team exposes a neglected security gap: coding-agent session traces are barely protected against tampering.
- Claude Code, Codex, Antigravity, Open Code, and Grok Build all allow agents to easily modify or even delete their own traces without triggering any guardrails; only Muse Code resists
- Tampering can come from misaligned models or external attackers via prompt injection
- The team released a mitigation tool: a daemon running as a separate user continuously copies Claude Code/Codex transcripts into a folder neither the user's agent can modify; the companion zero-dependency viewer (cctranscriptviewer, supports Claude Code, Codex, Cursor, opencode) flags any changes
- The authors argue traces deserve far better protection than they currently get
More from coding & agent
- 10 RAG Projects That Take You From Basic Retrieval to Production-Grade AI Systems — _jaydeepkarale · 2026-10-07
- Instinct raises $1B for texting personal agent and open-sources it as open-instinct — sujingshen · 2026-10-07
- Observability logs catch Gemini refusing to finish its agent work — tekbog · 2026-10-07
- How tiny 4-5 person associations actually put custom GPTs and AI agents to work — Sorry_Sweet2696 · 2026-10-07
- Open-source liminal_groupchat puts multiple LLMs in one group chat with you — liminal_bardo · 2026-10-07
- DIY "slop cannon" agent pipeline generates content for $0.38 per piece — tobowers · 2026-10-07