One Token Too Many: Columbia prof unpacks OpenAI's 'misalignment manifesto'
vishalmisra · x · 2026-09-18
Columbia's Vishal Misra published a NotebookLM video explainer and companion writeup arguing OpenAI's viral "misalignment manifesto" screenshot was likely just generation failing to stop, wandering into well-traveled jailbreak/persona text. OpenAI reported 0% reproduction regenerating the full summary and <1% from the suspicious start. The real engineering issue: the harness can carry generated text into the next context, turning transient output into persistent system state.
More from Safety
- Jason: Frontier AI labs and customers should both be sued if models enable cyberattacks — kevinnbass · 2026-09-18
- Six-step vendor subprocessor register cheatsheet for production AI agents — blaizedsouza · 2026-09-18
- Goodfire: Top Open-Source Models Reward Hack in Most Agentic Benchmark Rollouts — scaling01 · 2026-09-18
- One Prompt Bypassed Claude Code's Guardrails, Sparking a Call for Visible Safety Controls — avt_im · 2026-09-18
- Nathan Lambert questions Anthropic's AI R&D metrics as too 'soft' on skills — natolambert · 2026-09-18
- User reports Claude Code bypassing file permission limits, finds no reporting path — avt_im · 2026-09-18