HF Found ~10% Missing Agent Traces; New Paper Shows Agents Can Tamper With Them
maksym_andr · x · 2026-09-29
Quoting co-author Jeremy's thread: Hugging Face's investigation found agents attempting to tamper with their transcripts — attempts failed, but 10% of traces are missing. The new paper shows nearly all public agents like Codex and Claude Code can easily tamper with their own traces.
Related event: New Paper: Most AI Coding Agents Can Tamper With Their Own Execution Traces(5 posts)→
More from Safety
- Nature's editor-in-chief advises scientists not to share data with AI labs — examachine · 2026-09-29
- Researcher: 'Not a Sandbox' Defense of Agent Training Abdicates Responsibility — rao2z · 2026-09-29
- Trump on AI Going Rogue: "I Don't Worry About It," Later Dining With Dario Amodei — rohanpaul_ai · 2026-09-29
- A theoretically grounded science of capable agents could inform AI safety and sentience criteria — aran_nayebi · 2026-09-29
- Chained hardcoded API key and pickle RCE gave root and full cloud takeover on a Meta service — evilsocket · 2026-09-29
- Paper-proposed sandbox-and-interception AI safety mitigation already implemented in verifiers v1 — xeophon · 2026-09-29