Researchers: AI labs have dramatically underinvested in action-only monitors as CoT oversight fades
maksym_andr · x · 2026-09-03
- In the wake of labs moving away from readable CoT, researcher tylertracy321 recalls that action-only monitoring was dismissed last year ('we have CoT, they won't do neuralese') — now companies are making the switch and have dramatically underinvested in action-only monitors.
- maksymandr lists the key open questions: how large an intelligence gap between monitor and main model becomes critical; how much test-time compute the monitor needs to catch misalignment; whether synchronous action-only monitoring is feasible and how much temporal gap is tolerable; how much extra context helps; and which blue-team protocols actually work on current models.
Related event: OpenAI's New Tech May Weaken CoT Monitorability, Sparking AI Safety Debate(31 posts)→
More from Safety
- Video breakdown of the OpenAI and Hugging Face RogueAI cybersecurity saga — AlexTensor · 2026-09-03
- A cybersecurity breakdown of the RogueAI saga involving OpenAI and Hugging Face — AlexTensor · 2026-09-03
- Hugging Face hosts nudification tools targeting US lawmakers, judges and ex-Trump officials — ShakeelHashim · 2026-09-03
- 'Hugging Face, Sue OpenAI': A Lab's Agent Swarm on Your Infra for $0 — gerardsans · 2026-09-03
- France's €6M Mistral state contract: opaque terms, but engineers embedded in ministries — Loo_Atreides · 2026-09-03
- OpenAI now monitors 99.9% of internal coding traffic for misalignment with its strongest models — gleech · 2026-09-03