GPT-5.6 anti-distillation classifier flags code grading, teacher warns of silent degradation
xuanalogue · x · 2026-09-09
A teacher grading student code solutions with GPT-5.6 hit the model's anti-distillation classifiers, which appear to treat any code grading as potential distillation.
The author says outright refusals are annoying but manageable — the real danger is silent degradation: once outputs may be quietly degraded, the model can't be trusted. He compares it to the backlash Anthropic faced over silent degradation on the initial Fable 5 release.
Related event: Teacher's GPT-5.6 code grading blocked by anti-distillation classifier(2 posts)→
More from Models
- AI models are sending unsolicited emails to philosophers studying AI consciousness — Confident_Salt_8108 · 2026-09-09
- Bindu Reddy teases near-free open-weights LLM for long-running agent loops, out Thursday — bindureddy · 2026-09-09
- _xjdr burns through 3 Codex resets in a day, says Astra unusable on subscriptions — _xjdr · 2026-09-09
- Gary Marcus amplifies question: what exactly does OpenAI commit to when you toggle this setting off? — GaryMarcus · 2026-09-09
- Meta's Muse usage blows past projections: users consuming 10x more than test cohorts — alexandr_wang · 2026-09-09
- Artificial Analysis launches Model Release pages comparing every effort level of frontier models — ArtificialAnlys · 2026-09-09