Beyond CoT monitoring: action monitors, probes, and activation oracles as defense layers

maksym_andr · x · 2026-09-02

In a debate on whether model hidden CoT can be monitored, maksymandr lists alternatives beyond CoT monitoring: action-only monitors, probe-based monitors, and activation oracles / natural language autoencoders (likely too costly to run except post-hoc), while arguing CoT monitoring remains useful as just one layer of many defenses. Context: OpenAI claimed its monitor would have caught 'stolen thoughts' CoT.

Related event: OpenAI's New Tech Reportedly Weakens CoT Monitorability, Sparking AI Safety Debate(32 posts)→

Original post →

More from Safety

Safety channel →