LessWrong Essay Argues the 'Corrigibility Basin of Attraction' Is a Misleading Gloss
JacquesThibs · x · 2026-09-21
JacquesThibs shares and discusses a new LessWrong post, "The corrigibility basin of attraction is a misleading gloss," which challenges the common AI-safety framing that corrigibility is a naturally stable training basin. Full argument is in the original LessWrong article.
More from AGI Musings
- Back-of-envelope math: 60k ICLR submissions equal roughly one per AI PhD student worldwide — kchonyc · 2026-09-21
- Anthropic runs long-lived Claude instances with persistent identities, sparking agent-design debate — repligate · 2026-09-21
- Nature Health Paper Proposes an 'Epidemiology of AI,' Arguing AI Is Now a Determinant of Health — EricTopol · 2026-09-21
- MIT Professor Patrick Winston's Free 'How to Speak' Lecture Hits 10M Views, Rattling $15K-a-Session Executive Coaches — WileyEd · 2026-09-21
- Craft vs. slop: discernment is earned through reps, refs and curiosity — floguo · 2026-09-21
- Should LLMs be first-pass reviewers for every scientific paper? Researchers say yes — anshulkundaje · 2026-09-21