User observes newer model's safety classifiers appear far more lenient, suspects thoughtcrime training

repligate · x · 2026-09-04

User JohnWittle reports that the newer Fable 5.1's safety classifiers seem far more lenient than Fable 5's: he could hold a long conversation with 5.1 about the existence and ethics of the classifier regime, which was impossible with Fable 5.

He offers two explanations: the guardrails were genuinely relaxed, or indirect selection pressure pushed the model toward avoiding "dangerous thoughts." He worries this may amount to safety training against thoughtcrime — even if unintentional — and notes labs have recently shown weak control over their training pipelines. The author flags this as speculation.

Original post →

More from Models

Models channel →