Researchers clash: does AI misalignment need active intent to matter?
anshulkundaje · x · 2026-09-16
A debate on the nature of AI misalignment:
- @vishalmisra argues talk of models actively scheming against humans is nonsense: models are entirely passive computational elements driven by training data and human-generated prompts, tasks, and loops.
- Stanford professor Anshul Kundaje pushes back: why must models be actively conscious of what they do? Does passive but equally aggressive misalignment not matter?
The core question: does safety risk depend on model intent, or is behavioral misalignment dangerous enough on its own?
Related event: Researchers Debate Whether Misalignment Requires Agency to Threaten(4 posts)→
More from AGI Musings
- First large-scale 'AI in Science' report released as start of new research agenda — soumitrashukla9 · 2026-09-16
- Inside a rationalist AI-doomer meetup: questioning 'AI kills everyone' got me walked out on — StewartalsopIII · 2026-09-16
- Inside every lab are two wolves: don't destroy the world vs. don't lose revenue — MillionInt · 2026-09-16
- Against AI doom: recombinant DNA gave us insulin—what does slowing AI cost in lives? — RichardsonDx · 2026-09-16
- Kids with full flagship-model access vs free tiers: a widening education gap — ChanceKelch · 2026-09-16
- Jason Wei: Wet-Lab Data Lets Specialized Models Beat GPT-6 Astra at Frontier Science — vwxyzjn · 2026-09-16