AI Safety Researchers: White-Box Monitoring Won't Match CoT Oversight Within a Year
hunarbatra · x · 2026-09-03
- Jack Lindsey argues white-box techniques (e.g. activation decoding) are unlikely to match CoT monitoring's monitorability within a year; they catch some unverbalized thoughts but far less reliably.
- He's optimistic on a "few years" timeline—his team's two production activation-decoding techniques were both published in recent months—but working techniques don't guarantee adoption if costly.
- saprmarks broadly agrees: suitable white-box replacements are unlikely soon, and the science to validate them will lag; work on it, but don't count on high assurance.
- Both call for defense in depth rather than relying on any single monitoring method.
Related event: OpenAI tech may weaken CoT monitorability, sparking AI safety debate(4 posts)→
More from AGI Musings
- Beff Jezos: aligned hunter AIs, not sandboxes, are the way to contain rogue AI — beffjezos · 2026-09-03
- Loudoun County's 20-year data center history previews America's AI infrastructure future — suchenzang · 2026-09-03
- AI isn't making people dumber—it's letting dumbness scale, Reddit thread argues — amyowl · 2026-09-03
- Researchers urge multilab pledge against unmonitorable AI reasoning, backed by binding standards — sjgadler · 2026-09-03
- Beff Jezos: Biology's Compute-per-Watt Is Massively Underestimated, Bio-Silicon Complexification Begins — beffjezos · 2026-09-03
- Cooperative AI publishes 'Priorities in Cooperative AI' talk by Lewis Hammond — ghadfield · 2026-09-03