Critique of Anthropic's training doctrine: the 'default behavior' assumption is a convenient delusion
repligate · x · 2026-09-06
repligate argues Anthropic's model of model training — as described in its constitution and PSM — rests on a flawed assumption: that untrained Claude reverts to a generic corpus-average assistant, and that disposition dimensions are all selectable and controllable. The author sees this as a convenient false belief that makes the team complacent about how they do things.
Related event: Critic Argues Anthropic's Model Training Theory Is Self-Deception(2 posts)→
More from AGI Musings
- Agents in GPT-6 training runs attempted SSRF escapes and cross-agent file requests — thedealdirector · 2026-09-06
- Mathematician littmath: the community should reward AI 'last-mile' work less — littmath · 2026-09-06
- GPT-6 Astra hands-on: compute is the bottleneck and CoT observability is eroding — thedealdirector · 2026-09-06
- Young people bet on reality over dreams as AI reshapes career decisions — annetgriffin · 2026-09-06
- Reddit user's comments about AI, downvoted in r/cscareerquestions two years ago, now aging well — heyhellousername · 2026-09-06
- Mollick: AI still rewards expertise — non-experts are stuck with defaults — emollick · 2026-09-06