Labs and safety firms both have incentives: repligate on alignment spin and the CoT superstition

liminal_bardo · x · 2026-09-05

repligate amplifies a critique of AI safety narratives: labs are incentivized to claim "with our newest thing we've solved all these problems," while prosaic-safety firms like Redwood are incentivized to find new terrible problems to monitor. The truth is likely that the problems have always existed, remain unsolved, and have simply scaled. The quoted post by @sdmat123 argues that treating chain-of-thought as a privileged signal is superstition: CoT tokens are architecturally ordinary output shaped by post-training, so a strongly instruction-following model complying when told not to emit reasoning — or skipping reasoning on tasks it can answer cold — is unsurprising.

Original post →

More from AGI Musings

AGI Musings channel →