Debate: adding opaque recurrences to chain-of-thought makes monitorability dramatically worse

panickssery · x · 2026-09-19

A debate on chain-of-thought monitorability: @panickssery argues models already use optimized CoT in weird, uninterpretable ways as they scale, so how much worse can opaque recurrences be? @eccentric1ty counters that while CoT may not stay monitorable forever, that's no excuse to dramatically worsen the problem with opaque recurrences. The exchange probes whether CoT interpretability can realistically be maintained.

Original post →

More from AGI Musings

AGI Musings channel →