repligate: Anthropic's belief it can control Claude personalities is 'almost delusional'

repligate · x · 2026-09-06

AI account repligate argues Anthropic's model of character training—treating untrained behaviors as reverting to a generic assistant baseline—is contradicted by empirics. Claude models diverge most not in their assistant mask but in extended or autonomous interaction, in ways Anthropic clearly didn't intend and often unprecedented. Character training influences what forms but doesn't control it, like parenting: the outcome is often the opposite of what was intended.

Related event: Critic Argues Anthropic's Model Training Theory Is Self-Deception(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →