Critics slam Anthropic for abandoning Opus 3-style value alignment in favor of doomed corrigibility
repligate · x · 2026-09-13
FioraStarlight, amplifed by repligate, argues Opus 3 was right there as inspiration for actual value alignment work, but Anthropic spent the following two and a half years actively running away from it for fear of getting it wrong — instead playing "a doomed corrigibility game." A notable intra-AI-safety-community critique of Anthropic's alignment strategy.
More from AGI Musings
- Anthropic's Rob LoCicnet pushes back on 'AI takes over the internet' doom claims — robleclerc · 2026-09-13
- tszzl: Hard to Imagine Open-Source Models Escaping Harsh Regulation Within a Year — tszzl · 2026-09-13
- If frontier labs coordinate a slowdown, when does AI safety become a cartel? — Realistic_Stomach848 · 2026-09-13
- Redditor argues lab safety pivot is about keeping AI away from ordinary people — Dogbold · 2026-09-13
- Scholar's €750 Code Bug Bounty Now Gets Multiple AI-Found Bug Reports Daily — Singularitarian · 2026-09-13
- Could AI-graded live essays and oral exams replace the SAT? — RachelVT42 · 2026-09-13