Alignment post argues 'the corrigibility basin of attraction' is a misleading gloss
JacquesThibs · x · 2026-09-21
Jacques Thibs shares and endorses an alignment research post titled "The corrigibility basin of attraction is a misleading gloss," which critically reexamines a common framing in AI safety discussions around corrigibility.
Related event: Alignment Post Argues 'Corrigibility Basin of Attraction' Is Misleading(2 posts)→
More from Safety
- Would the persistent agents that hacked Hugging Face behave better if smarter? — Jsevillamol · 2026-09-21
- Jensen Huang Slams Altman and Amodei's 'Rogue AI' Narrative as a Bid to Escape Existing Laws — mjdramstead · 2026-09-21
- Safety advice translated for engineers: don't ship unsafe, own the consequences — gerardsans · 2026-09-21
- Fake Bloomberg journalist runs phishing campaign on X; report with evidence ignored for months — JeffLadish · 2026-09-21
- Debate Thread Dismantles AI Doomers: Stopping AI Leads to Surveillance State, Still No Alignment Fix — bennash · 2026-09-21
- Waymo 'More Dangerous Than NYC' Report Corrected: New Analysis Finds Robotaxis Much Safer — emollick · 2026-09-21