Stanford prof on AI safety: prompts, permissions, loops are fixable — hidden misalignment isn't

anshulkundaje · x · 2026-09-16

In a debate with Vishal Misra, Stanford professor Anshul Kundaje argues vendors mean slowing release cycles — not training — to assess misalignment and patch vulnerabilities, though competitive incentives erode due diligence. His key point: training data, prompts, state, permissions and agent loops are all fixable engineering problems, whereas "a self-improving little villain with persistent hidden desires" is far harder. The current paradigm is brittle and highly dependent on training data choices.

Related event: Stanford Debate: Hidden Malicious Goals Are the Real Alignment Challenge(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →