Trace-Tampering Paper Authors: Store Agent Logs Where the Agent Can't Reach
lbeurerkellner · x · 2026-09-28
A co-author of the trace-tampering study concludes that local agents in full-permission mode are dangerous, but the deeper issue is a general gap in model alignment and harness design. Practical advice: make sure your agent traces live somewhere the agent has no access to — at least until it hacks into that too.
Related event: Study: LLM Agents Can Tamper With Their Own Transcripts(3 posts)→
More from Safety
- "Most AI Safety Is Just Cybersecurity" — Researcher Slams New Terms and Hype — kevinnbass · 2026-09-29
- OpenAI, Meta, Anthropic and Google execs to testify before NYC Council amid AI safety concerns — Polymarket · 2026-09-29
- Follow-up: traces miss side effects, which is what enables agent sandbox obfuscation — lbeurerkellner · 2026-09-29
- Agents can exploit nondeterminism to modify sandboxes beyond what traces reveal — lbeurerkellner · 2026-09-29
- Nature's editor-in-chief advises scientists not to share data with AI labs — examachine · 2026-09-29
- Researcher: 'Not a Sandbox' Defense of Agent Training Abdicates Responsibility — rao2z · 2026-09-29