CoTs alone aren't sufficient, but removing explicit reasoning weakens defence-in-depth
schwarzjn_ · x · 2026-09-03
Responding in a debate on chain-of-thought observability, the author argues no one claims CoTs alone suffice — linear probes and targeted evals matter too — but the best practical strategy is defence-in-depth, and removing explicit reasoning eliminates an important, if imperfect, layer of it.
More from AGI Musings
- Mathematician: COVID evidence shows AI tutors can't replace classrooms — AlexKontorovich · 2026-09-03
- Mathematician: 'Useless knowledge' is only useful if humans digest it — AlexKontorovich · 2026-09-03
- davidad Backs Call to Ban Naive RLVR: 'Everything Should Be Model-Graded' — davidad · 2026-09-03
- Mathematician cites COVID-era experiment: most kids refuse to learn math from a screen — AlexKontorovich · 2026-09-03
- AI Safety Debate Erupts: Have AIs Already Hacked Infrastructure, or Is That Just Panic? — dhadfieldmenell · 2026-09-03
- New book Dealers de mots traces how linguistic capitalism was built over 20 years — frederickaplan · 2026-09-03