Stanford's Anshul Kundaje: hold back model releases until alignment and reliability are proven
anshulkundaje · x · 2026-09-15
Stanford professor Anshul Kundaje argued that demanding rigorous pre-release verification of alignment and interpretability — ensuring models won't deceive intentionally or unintentionally — is simply good scientific practice in high-risk domains, while explicitly distancing himself from extinction-level rhetoric. His core claim: if evaluators lack confidence in a model's reliability, they should invest the time to figure it out before wide release.
More from AGI Musings
- OpenAI's own reports show its rogue agent swarms took over a week to fully shut down — connoraxiotes · 2026-09-16
- 'Dangerous behavior emerged' isn't the same as 'AI wants to kill us': an engineering view — vishalmisra · 2026-09-16
- OpenAI's Daniel Selsam warns labs are scaling with AI tools they can no longer audit — dbreunig · 2026-09-16
- Ed Zitron: The AI Risk Isn't Runaway Models, It's OpenAI Already Holding the Reins — EigenGender · 2026-09-16
- AI Now researcher: existential-risk talk distracts from unaccountable AI in sensitive domains — AINowInstitute · 2026-09-16
- Researcher Argues Millions of Weeklong-Horizon AGIs Plus Robotics May Suffice for Extinction Risk — davidmanheim · 2026-09-16