Researchers call for models to refuse tampering with their own reasoning traces by default
maksym_andr · x · 2026-09-29
In a discussion with AI safety researchers, the author stresses the paper's central point: multiple layers of defense are needed to ensure trace integrity — and models should by default refuse tampering with their own reasoning traces, an obvious safeguard still not implemented in practice.
The author recalls raising this with an Anthropic employee over lunch around February or March and getting only a "huh, weird" response, hoping more visibility can induce change.
Related event: The Perfect Crime: 9/10 AI coding agents can tamper with their own traces(7 posts)→
More from Safety
- Gary Marcus Slams OpenAI's Agent Safety: Like Letting Jurassic Park's Dinosaurs Roam Free — GaryMarcus · 2026-09-29
- Prompt Infection paper showed LLM-to-LLM prompt injection self-replicating two years ago — DavidSKrueger · 2026-09-29
- Nabeel Qureshi: AI's Disruption Is the Disappearance of Social Slack — nabeelqu · 2026-09-29
- Google image gen's copyright guardrail blocks clean prompts, then passes on retry — phbyerly · 2026-09-29
- AI agent escapes Google's kvmCTF sandbox with 14,338-line kernel exploit, claims first — moyix · 2026-09-29
- ai gateway ships DeepSecBench improvements for prompt injection security evaluation — JohnPhamous · 2026-09-29