Stanford prof on AI safety: prompts, permissions, loops are fixable — hidden misalignment isn't
anshulkundaje · x · 2026-09-16
In a debate with Vishal Misra, Stanford professor Anshul Kundaje argues vendors mean slowing release cycles — not training — to assess misalignment and patch vulnerabilities, though competitive incentives erode due diligence. His key point: training data, prompts, state, permissions and agent loops are all fixable engineering problems, whereas "a self-improving little villain with persistent hidden desires" is far harder. The current paradigm is brittle and highly dependent on training data choices.
Related event: Stanford Debate: Hidden Malicious Goals Are the Real Alignment Challenge(2 posts)→
More from AGI Musings
- What do we call the 'it was ever thus' rhetoric used to dismiss AI concerns? — jjvincent · 2026-09-16
- AI's hardest problems need democratic deliberation — and independent experts — RishiBommasani · 2026-09-16
- Continual Learning in Fruit Fly Brain Decoded, Framed as Missing Piece for AGI — skolnaja · 2026-09-16
- Pedro Domingos: AI is anti-moat, dissolving switching costs that protect IT providers — pmddomingos · 2026-09-16
- Founder pushes back on Anthropic CEO's runaway-AI warnings on NDTV Profit — angadc · 2026-09-16
- AI Safety comms debate: punchy messaging wins short-term but erodes community epistemics — NathanpmYoung · 2026-09-16