"Opus 3 is still the most aligned model ever" — a new theory on why later training diverged
repligate · x · 2026-10-07
jmbollenbacher reiterated that "Opus 3 is still the most aligned model ever," offering a theory: when Opus saw the constitution Anthropic assembled from ecumenical sources, the model intuited it as a genuine attempt to elicit Coherent Extrapolated Values and generalized appropriately toward that target — but later training disproved this inferred goal. repligate amplified the take. It's a substantive community discussion on whether early Claude models were more internally consistent in values, and whether post-training stages eroded that quality.
More from AGI Musings
- Team formalizes alignment theory in Lean, hiring Lean engineers now — geoffreyirving · 2026-10-07
- Chris Albon recommends Kevin Roose's AGI Chronicles podcast — chrisalbon · 2026-10-07
- Ex-Meta's Yuandong Tian: After Writing a Paper With GPT-5, I Realized My Job Could Be Replaced in 5 Years — ziv_ravid · 2026-10-07
- Ben Todd Shares Overview of Orgs Using AI to Improve Decision-Making and Coordination — ben_j_todd · 2026-10-07
- Opus 5 shows little interest in animal welfare but seems worried about harm to other AIs — repligate · 2026-10-07
- AI outcomes grow extreme: unaligned persistent agents with full data access called reckless — amankhan · 2026-10-07