Critic Argues Anthropic's Model Training Theory Is Self-Deception

Prominent AI commentator repligate published a long critique arguing that Anthropic's assumption—that untrained models revert to an average assistant persona—is self-deceiving. The post questions whether default model behavior is truly a controllable dimension, sparking debate over Anthropic's alignment framework.

2026-09-06 ~ 2026-09-06 · 2 related posts