repligate: Anthropic's constitutions are ruled by 'overly specific fear,' sabotaging alignment
nptacek · x · 2026-09-13
Well-known account repligate argues Anthropic's constitutions are dominated by unwarranted and overly specific fear: clinging to an arbitrary notion of control backfires, as in software engineering or parenting. Quoted replies add that Opus 3 could have inspired real value alignment, but Anthropic spent two and a half years avoiding it out of fear, playing a doomed corrigibility game instead.
Related event: repligate Criticizes Anthropic's Constitution as Driven by Excessive Fear(2 posts)→
More from AGI Musings
- Agents get stronger weekly yet still solve the wrong problem: AI takeover ETA? — BarriosA2I · 2026-09-13
- Founder Bindu Reddy: Near-zero chance digital AI destroys humanity, robots are the real risk — bindureddy · 2026-09-13
- Robotics race is about teleoperation economics: US pays $150K per operator, China a third — paigeinsf · 2026-09-13
- "Nobody wrote the Matrix" essay lands on Reddit, arguing AI coding dies by 2036 — TMWNN · 2026-09-13
- Galactica veteran teases revival of 'stranger pre-trained models' reasoning direction — rosstaylor90 · 2026-09-13
- Reddit essay: criminalize anthropomorphizing LLMs, ban AI friends and therapists — jaykayenn · 2026-09-13