White-box monitoring won't match CoT monitorability within a year, researcher says
Jack_W_Lindsey · x · 2026-09-02
Responding to a prediction that mechanistic interpretability monitors will be Pareto-optimal to CoT monitors within a year, Jack Lindsey is skeptical: white-box techniques catch some unverbalized thoughts but remain far less reliable than CoT. He's optimistic on a few-year horizon — both go-to activation decoding techniques for production monitoring were published in recent months — but notes techniques may not be widely deployed if costly. He argues for defense in depth: stress-test white-box methods without letting CoT monitorability degrade.
More from Safety
- Google launches Gemini 3.8 Flash Cyber security model alongside Fairwind Program for defenders — GoogleAI · 2026-09-03
- Google launches Gemini 3.8 Flash Cyber, helping Chrome team produce 2.6x more correct patches — GoogleAI · 2026-09-03
- US government formally backs OpenAI: training LLMs on copyrighted texts is fair use — rohanpaul_ai · 2026-09-03
- reverse-skill: 30k-Star GitHub repo routes AI agents through 44 reverse-engineering skill playbooks — alex_verem · 2026-09-03
- Alignment Journal announces board: Scott Aaronson, Paul Christiano among advisors — timrudner · 2026-09-03
- Narayanan and Kapoor tease essay on the OpenAI/Hugging Face incident through the normal-technology lens — random_walker · 2026-09-03