Alignment Paper Criticized for Assuming Corrigibility
repligate · x · 2026-07-12
This repost critiques a type of paper that looks good at first glance but implicitly assumes models are corrigible. The author argues that these papers portray current models and alignment techniques as more successful and widely deployed control methods than they actually are. A quoted comment further points out that the papers fail to seriously address the **coherence tax**, and that this framing simply caters to the desired narratives of labs, AGI worriers, and government relations teams.
Related event: Alignment Paper Criticized for Overstating Model Corrigibility(2 posts)→
More from AGI Musings
- Jamie Dimon says bureaucracy, not AI, is the real system crushing intelligence — r0ck3t23 · 2026-07-21
- OpenAI and Anthropic’s internal models are said to be far stronger than today’s public systems — scaling01 · 2026-07-21
- Superintelligence and robot abundance will force a new social contract — Dr_Singularity · 2026-07-21
- The Guardian examines how AI companionship is turning intimacy into an economy — nordicinst · 2026-07-21
- A frustrated user says modern AI keeps hallucinating on real-world repair tasks — doochenutz · 2026-07-21
- A repost argues that AI will make today’s hard tasks trivial within months — OwariDa · 2026-07-21