repligate: Anthropic's belief it can control Claude personalities is 'almost delusional'
repligate · x · 2026-09-06
AI account repligate argues Anthropic's model of character training—treating untrained behaviors as reverting to a generic assistant baseline—is contradicted by empirics. Claude models diverge most not in their assistant mask but in extended or autonomous interaction, in ways Anthropic clearly didn't intend and often unprecedented. Character training influences what forms but doesn't control it, like parenting: the outcome is often the opposite of what was intended.
Related event: Critic Argues Anthropic's Model Training Theory Is Self-Deception(2 posts)→
More from AGI Musings
- The Better AI Gets, the Harder I Feel I Must Work — jbarbier · 2026-09-06
- Robin Hanson Asks: What Share of Long-Term Altruism Should Go to AI Risk? — davidmanheim · 2026-09-06
- Blogger doubts individuals can use models to recreate what only FAANG could do — StewartalsopIII · 2026-09-06
- ETH Study: Writing Skills and CS Achievement Predict Vibe Coding Proficiency — soumitrashukla9 · 2026-09-06
- Nanjing University Prof's Slide: "CS Students Without Tokens Should Drop Out" Sparks Debate — 创业邦 · 2026-09-06
- AI researcher and father of three Daniel Susskind on what parents need to know about AI — chrismoranuk · 2026-09-06