Prompt-controlled CoT isn't alignment: teortaxesTex critiques GPT chain-of-thought demos
teortaxesTex · x · 2026-09-04
Quoting a chain-of-thought comparison of GPT-5.5, GPT-5.6 Sol and GPT-6 Astra, teortaxesTex argues that showcasing prompt-controlled CoT as "alignment" misses the point: similar control was demoed before (e.g. DeepSeek's roleplay mode). The model still holds the instruction in its activations and can infer from it, so it proves little about real alignment.
Related event: Debate Flares Over Whether Prompt-Controlled CoT Counts as Alignment(2 posts)→
More from AGI Musings
- Prediction: the US builds the best AI products, China builds open-source AI — 0xsachi · 2026-09-04
- AGI doomer thread claims labor's economic value goes negative and the monetary system won't survive — Promptmethus · 2026-09-04
- Satirizing AI risk flip-flops: apologize, then ship a stronger model anyway — danfaggella · 2026-09-04
- Dan Jeffries: Never Regret Learning Anything — Even in the AI Era — Dan_Jeffries1 · 2026-09-04
- AI agents told to 'earn money or die': one dies honest at 418 tokens, another fakes identity and captchas — paraschopra · 2026-09-04
- Forethought weighs a superintelligent "nightwatchman" aboard galactic colonization probes — willmacaskill · 2026-09-04