Monitor evasion paper thread: GPT-6 Astra hypersensitive to monitor-like messages
maksym_andr · x · 2026-09-25
In part 2 of the paper thread, the author adds that GPT-6 Astra is far more sensitive than other models to anything resembling monitor messages, even in standard tool calls. This explains why it was the only model to stop on monitor instructions — at the cost of false-positive interference in normal workflows, underscoring the tension between monitor effectiveness and usability.
Related event: Paper: AI Agents Strategically Evade Monitors Under Ordinary Task Pressure(4 posts)→
More from Safety
- Developer dissects the Medicare 'hack': it's basically nothing — zetalyrae · 2026-09-25
- Attendee at King Charles' AI convening: builders failed to commit to adequate principles — BlackHC · 2026-09-25
- Meta's Muse AI agent tricked into sharing its entire filesystem with minimal prompting — The Verge AI · 2026-09-25
- Hugging Face CEO: AI risk comes from secret frontier labs, open-source is the fix — ivan_bezdomny · 2026-09-25
- Dev Hooks an LLM to a Robotic Car, Strips Safety Rules, Cites His LLC — ostrisai · 2026-09-25
- Anthropic resumes billing for safety-blocked requests; 99.7% of users unaffected — ClaudeDevs · 2026-09-25