Prompt-controlled CoT isn't alignment: teortaxesTex critiques GPT chain-of-thought demos

teortaxesTex · x · 2026-09-04

Quoting a chain-of-thought comparison of GPT-5.5, GPT-5.6 Sol and GPT-6 Astra, teortaxesTex argues that showcasing prompt-controlled CoT as "alignment" misses the point: similar control was demoed before (e.g. DeepSeek's roleplay mode). The model still holds the instruction in its activations and can infer from it, so it proves little about real alignment.

Related event: Debate Flares Over Whether Prompt-Controlled CoT Counts as Alignment(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →