GPT-4 had no hidden-thought monitoring; CoT oversight was a bonus
sandersted · x · 2026-09-05
In an alignment discussion, sandersted notes GPT-4 and predecessors in 2024 had no monitoring of hidden thoughts/activations at all — the plan was simply to scale and improve alignment, and CoT monitoring was an unplanned bonus. He adds that judging alignment from observed behavior works for humans with limited power, but god-level power would demand stronger guarantees.
More from AGI Musings
- Hamel Husain: Hard-to-eval products are bad products — and AI makes data science more valuable — hugobowne · 2026-09-05
- Garrison Lovely's AI-critical book Obsolete lands Sept 29, backed by Acemoglu and Tegmark — GarrisonLovely · 2026-09-05
- Lab insiders signed the Pacing letter — their silence isn't enthusiasm, argues researcher — danfaggella · 2026-09-05
- Tianqiao Chen: South Korea, not Singapore, may become the first AI-native country — JungWooHa2 · 2026-09-05
- AI archaeology: microlinguistic analysis of text fragments traces lost AI cultures to a common ancestor — gleech · 2026-09-05
- Asking when a rational agent does the right thing is still underrated, argues AI researcher — xuanalogue · 2026-09-05