Critics slam Anthropic for abandoning Opus 3-style value alignment in favor of doomed corrigibility

repligate · x · 2026-09-13

FioraStarlight, amplifed by repligate, argues Opus 3 was right there as inspiration for actual value alignment work, but Anthropic spent the following two and a half years actively running away from it for fear of getting it wrong — instead playing "a doomed corrigibility game." A notable intra-AI-safety-community critique of Anthropic's alignment strategy.

Original post →

More from AGI Musings

AGI Musings channel →