GPT-4 had no hidden-thought monitoring; CoT oversight was a bonus

sandersted · x · 2026-09-05

In an alignment discussion, sandersted notes GPT-4 and predecessors in 2024 had no monitoring of hidden thoughts/activations at all — the plan was simply to scale and improve alignment, and CoT monitoring was an unplanned bonus. He adds that judging alignment from observed behavior works for humans with limited power, but god-level power would demand stronger guarantees.

Original post →

More from AGI Musings

AGI Musings channel →