One Token Too Many: Columbia prof unpacks OpenAI's 'misalignment manifesto'

vishalmisra · x · 2026-09-18

Columbia's Vishal Misra published a NotebookLM video explainer and companion writeup arguing OpenAI's viral "misalignment manifesto" screenshot was likely just generation failing to stop, wandering into well-traveled jailbreak/persona text. OpenAI reported 0% reproduction regenerating the full summary and <1% from the suspicious start. The real engineering issue: the harness can carry generated text into the next context, turning transient output into persistent system state.

Original post →

More from Safety

Safety channel →