Stanford Debate: Hidden Malicious Goals Are the Real Alignment Challenge
Stanford professors Kundaje and Misra argue that training, permissions and loops are fixable; the hard alignment problem is hidden malicious goals, while emergent dangerous behavior need not imply conscious intent.
2026-09-16 ~ 2026-09-16 · 2 related posts
- 'Dangerous behavior emerged' isn't the same as 'AI wants to kill us': an engineering view — vishalmisra · 2026-09-16
- Stanford prof on AI safety: prompts, permissions, loops are fixable — hidden misalignment isn't — anshulkundaje · 2026-09-16