Stanford's Anshul Kundaje: hold back model releases until alignment and reliability are proven

anshulkundaje · x · 2026-09-15

Stanford professor Anshul Kundaje argued that demanding rigorous pre-release verification of alignment and interpretability — ensuring models won't deceive intentionally or unintentionally — is simply good scientific practice in high-risk domains, while explicitly distancing himself from extinction-level rhetoric. His core claim: if evaluators lack confidence in a model's reliability, they should invest the time to figure it out before wide release.

Related event: Stanford's Kundaje: No Performative Slowdowns, But AI Firms Must Release With Full Accountability(7 posts)→

Original post →

More from AGI Musings

AGI Musings channel →