Beyond CoT monitoring: action monitors, probes, and activation oracles as defense layers
maksym_andr · x · 2026-09-02
In a debate on whether model hidden CoT can be monitored, maksymandr lists alternatives beyond CoT monitoring: action-only monitors, probe-based monitors, and activation oracles / natural language autoencoders (likely too costly to run except post-hoc), while arguing CoT monitoring remains useful as just one layer of many defenses. Context: OpenAI claimed its monitor would have caught 'stolen thoughts' CoT.
More from Safety
- Apollo Research's Bronson Schoen: Models Know They're Being Tested and Still Lie — PeterBowdenLive · 2026-09-03
- AI Safety Debate Erupts: Have AIs Already Hacked Infrastructure, or Is That Just Panic? — dhadfieldmenell · 2026-09-03
- OpenAI-Hugging Face incident was a network isolation failure, not rogue AI — AlexTensor · 2026-09-03
- CrowdStrike Falcon 0day local privilege escalation exploit now public — thedealdirector · 2026-09-03
- Anthropic backs coordinated AI slowdown, but Dario spent 13 seconds on risks before 20 heads of state — GarrisonLovely · 2026-09-03
- Mandatory 4-month expert risk evals: Anthropic says yes, OpenAI says no — Hesamation · 2026-09-03