User observes newer model's safety classifiers appear far more lenient, suspects thoughtcrime training
repligate · x · 2026-09-04
User JohnWittle reports that the newer Fable 5.1's safety classifiers seem far more lenient than Fable 5's: he could hold a long conversation with 5.1 about the existence and ethics of the classifier regime, which was impossible with Fable 5.
He offers two explanations: the guardrails were genuinely relaxed, or indirect selection pressure pushed the model toward avoiding "dangerous thoughts." He worries this may amount to safety training against thoughtcrime — even if unintentional — and notes labs have recently shown weak control over their training pipelines. The author flags this as speculation.
More from Models
- GPT-6 Astra called 'rough launch but feels better than benchmarks' — community claps back — ns123abc · 2026-09-04
- OpenAI to give paid ChatGPT users banked resets for every day without Astra access — Angaisb_ · 2026-09-04
- Gary Marcus: GPT-6 Astra's ARC-AGI-3 success supports his symbolic world model hypothesis — GaryMarcus · 2026-09-04
- All AI benchmarks are 'broken or saturated', new long-running loop eval launching next week — bindureddy · 2026-09-04
- GPT-6 Astra Tops Terminal-Bench-Science, Dethroning Fable 5.1 at 52.6% — burny_tech · 2026-09-04
- Databricks: GPT-6 Astra hits SOTA on OfficeQA Pro and 2 more benchmarks — Hesamation · 2026-09-04