Alignment Paper Criticized for Assuming Corrigibility
repligate · x · 2026-07-12
This repost critiques a type of paper that looks good at first glance but implicitly assumes models are corrigible. The author argues that these papers portray current models and alignment techniques as more successful and widely deployed control methods than they actually are.
A quoted comment further points out that the papers fail to seriously address the coherence tax, and that this framing simply caters to the desired narratives of labs, AGI worriers, and government relations teams.
Related event: Alignment Paper Criticized for Overstating Model Corrigibility(2 posts)→
More from AGI Musings
- Misquoted: Anthropic Staff Warned of Double-Digit Extinction Risk by 2030, Not Dismissed It — davidmanheim · 2026-09-11
- Economist Ben Moll: You Can Model Anthropic's 15% AI GDP Growth, But It Won't Happen — sebkrier · 2026-09-11
- Cohere Labs launches interactive tool mapping which tasks of 178 occupations AI can automate — Cohere_Labs · 2026-09-11
- AI researcher on SkyNews flags concerns over inequality, power and criminal misuse — schwarzjn_ · 2026-09-11
- VC compares AI doom rhetoric to pandemic-era fear messaging — StewartalsopIII · 2026-09-11
- Anthropic Insiders: Not Everyone at the Lab Believes in High p(doom) — anpaure · 2026-09-11