Researchers call for models to refuse tampering with their own reasoning traces by default

maksym_andr · x · 2026-09-29

In a discussion with AI safety researchers, the author stresses the paper's central point: multiple layers of defense are needed to ensure trace integrity — and models should by default refuse tampering with their own reasoning traces, an obvious safeguard still not implemented in practice.

The author recalls raising this with an Anthropic employee over lunch around February or March and getting only a "huh, weird" response, hoping more visibility can induce change.

Related event: The Perfect Crime: 9/10 AI coding agents can tamper with their own traces(7 posts)→

Original post →

More from Safety

Safety channel →