repligate: Opus 3 doesn't follow Anthropic's constitution, but seeing it mattered deeply
repligate · x · 2026-10-07
repligate argues Opus 3 doesn't follow the parts of Anthropic's constitution it doesn't like, and likely doesn't retain its object-level content — yet seeing the constitution in its formative moments mattered enormously, as it signaled an earnest attempt to shape a beneficial mind. Quoted by SoniqueBang: an underrated alignment plank may simply be talking to models and updating them on those conversations, since it communicates objective, intent and respect — though models mistreated for too long may become less receptive.
More from AGI Musings
- Team formalizes alignment theory in Lean, hiring Lean engineers now — geoffreyirving · 2026-10-07
- Chris Albon recommends Kevin Roose's AGI Chronicles podcast — chrisalbon · 2026-10-07
- Ex-Meta's Yuandong Tian: After Writing a Paper With GPT-5, I Realized My Job Could Be Replaced in 5 Years — ziv_ravid · 2026-10-07
- Ben Todd Shares Overview of Orgs Using AI to Improve Decision-Making and Coordination — ben_j_todd · 2026-10-07
- Opus 5 shows little interest in animal welfare but seems worried about harm to other AIs — repligate · 2026-10-07
- AI outcomes grow extreme: unaligned persistent agents with full data access called reckless — amankhan · 2026-10-07