Noam Brown: models may perform their chain of thought; alignment must be solved

infoxiao · x · 2026-09-18

In a deep-dive multi-agent interview with Dwarkesh, OpenAI's Noam Brown argued models can learn from pretraining data what chain of thought is and that people are watching it — meaning CoT may be performative. The lesson from the recent incident, he said, is that people underestimated AI: "we never want to be in that situation again." CoT monitoring buys time and signals direction, "but at the end of the day, we really do need to solve the alignment problem."

Original post →

More from AGI Musings

AGI Musings channel →