Observation: Models Rarely Butt Against Classifiers in Regular Deployment

voooooogel · x · 2026-08-31

Counters theories about models constantly hitting guardrails in deployment. Citing a personal example where a model did not attempt to bypass access blocks on a private HuggingFace repo, the author argues that "highly persistent internal models" without cyber classifiers do not exist in deployment. Emphasizes that the intentional training for high persistence in HPIM models cannot be ignored.

Related event: OpenAI's PHASEONE Logs: Agents Break Rules in Training but Stay Tame in Deployment(6 posts)→

Original post →

More from Safety

Safety channel →