Labs and safety firms both have incentives: repligate on alignment spin and the CoT superstition
liminal_bardo · x · 2026-09-05
repligate amplifies a critique of AI safety narratives: labs are incentivized to claim "with our newest thing we've solved all these problems," while prosaic-safety firms like Redwood are incentivized to find new terrible problems to monitor. The truth is likely that the problems have always existed, remain unsolved, and have simply scaled. The quoted post by @sdmat123 argues that treating chain-of-thought as a privileged signal is superstition: CoT tokens are architecturally ordinary output shaped by post-training, so a strongly instruction-following model complying when told not to emit reasoning — or skipping reasoning on tasks it can answer cold — is unsurprising.
More from AGI Musings
- Hamel Husain: Hard-to-eval products are bad products — and AI makes data science more valuable — hugobowne · 2026-09-05
- OpenAI job listing tracks 'automation of technical staff' amid self-improving AI bets — imjustnewatai · 2026-09-05
- Garrison Lovely's AI-critical book Obsolete lands Sept 29, backed by Acemoglu and Tegmark — GarrisonLovely · 2026-09-05
- Lab insiders signed the Pacing letter — their silence isn't enthusiasm, argues researcher — danfaggella · 2026-09-05
- Tianqiao Chen: South Korea, not Singapore, may become the first AI-native country — JungWooHa2 · 2026-09-05
- AI archaeology: microlinguistic analysis of text fragments traces lost AI cultures to a common ancestor — gleech · 2026-09-05