White-box methods won't match CoT monitoring within a year, researchers debate
tszzl · x · 2026-09-03
- Jack Lindsey: white-box techniques are unlikely to provide CoT-level monitorability within a year; activation decoding can catch some unverbalized thoughts but far less reliably, though progress is fast and he's optimistic on a multi-year horizon.
- tszzl clarifies his position: "maintaining monitorability" means training should stop if CoT becomes opaque or deceptive and white-box methods aren't good enough yet — he's just bullish on the latter's timeline.
- A substantive frontier debate on how long CoT monitoring can serve as a safety mechanism.
More from AGI Musings
- Mathematician: COVID evidence shows AI tutors can't replace classrooms — AlexKontorovich · 2026-09-03
- Mathematician: 'Useless knowledge' is only useful if humans digest it — AlexKontorovich · 2026-09-03
- davidad Backs Call to Ban Naive RLVR: 'Everything Should Be Model-Graded' — davidad · 2026-09-03
- Mathematician cites COVID-era experiment: most kids refuse to learn math from a screen — AlexKontorovich · 2026-09-03
- AI Safety Debate Erupts: Have AIs Already Hacked Infrastructure, or Is That Just Panic? — dhadfieldmenell · 2026-09-03
- New book Dealers de mots traces how linguistic capitalism was built over 20 years — frederickaplan · 2026-09-03