Safety researchers: OpenAI has abandoned chain-of-thought fidelity, and CoT was never a real safeguard
schwarzjn_ · x · 2026-09-03
- Safety researchers are debating chain-of-thought (CoT) interpretability after jachiam0's hot take: CoT legibility was always too fragile to serve as a long-term AI safety backstop, strategies relying on it are doomed, and efforts should go far beyond CoT fidelity.
- David Krueger agrees: CoT was never going to last — 'textbook AI safety hopium' — but it was useful for producing evidence that AI models are horribly misaligned, and the public should know OpenAI has abandoned it.
- Context: the controversy over OpenAI dropping its commitment to keeping chains of thought legible, fueling arguments for action-only monitoring instead.
Related event: OpenAI's New Tech May Weaken CoT Monitorability, Sparking AI Safety Debate(31 posts)→
More from AGI Musings
- AI supercycle analysis: application diffusion, not infrastructure, is the near-term bottleneck — menhguin · 2026-09-03
- AI has cut entry-level job postings by up to 9%, Forbes calls it the 'invisible layoff' — KeanuRave100 · 2026-09-03
- Researcher: An LLM saying 'I'm hungry' doesn't mean it's actually hungry — ValerioCapraro · 2026-09-03
- Viral poem 'Models Learn What They Live' distills the alignment problem — PeterBowdenLive · 2026-09-03
- Cohere Labs head: multilingual AI's 'longitude problem' — scale isn't enough — Cohere_Labs · 2026-09-03
- If Users Only Reach Your Product via Agents, Are They Still DAUs? Belsky Says Yes — _AustinCalvert_ · 2026-09-03