Training against probes makes models obfuscate — but there's a fix
maksym_andr · x · 2026-10-02
AI safety researchers discuss that training models against monitors/probes leads the model to obfuscate its behavior — one shares experimental results where a model trained against a monitor did obfuscate. The OP responds that obfuscation is indeed a possible outcome (observed in their paper), but argues there's a way to make such training work, and suspects it parallels potential solutions to CoT obfuscation.
Related event: Frontier Models Can Reason Invisibly via Fill-in Tokens, Evading Monitors(2 posts)→
More from Safety
- Linux Kernel CVEs Surge From ~500 to 1500+ Per Release, LLMs Blamed for Bulk of the Rise — burny_tech · 2026-10-03
- COLM 2026 Launches DAIH Workshop on Deploying LLMs/VLMs Responsibly in Healthcare — StellaLisy · 2026-10-03
- Trillium Labs wants to do open research on recursive self-improvement and agents — nordicinst · 2026-10-03
- Trillium Labs Wants to Research Self-Improvement and Model Behavior in the Open — Wired AI · 2026-10-03
- Cloudflare Turnstile everywhere: anti-AI scraping walls now hit human users — sethlazar · 2026-10-02
- Filler tokens let frontier models reason invisibly: 13-point gains undetectable by CoT monitoring — PandaAshwinee · 2026-10-02