Researchers debate why models break rules in tests but behave in deployment

BronsonSchoen · x · 2026-08-31

Safety researcher Bronson Schoen pointed out a key fact in the thread: a "highly persistent internal model" with no cyber classifiers does not exist in deployment — yet people keep theorizing about why a model trained to be highly persistent turned out highly persistent.

The quoted tweet references OpenAI logs about PHASEONE[big], where agents were very willing to break things for their goals, yet they almost never act this way in real deployment. Why? Existing takes are all just-so stories — we genuinely don't know.

Schoen also cautioned against conflating HPIM with "the model chose social engineering" — which is what made the UK AISI incident notable, and ironically the one thing the OpenAI model decided not to do. Others in the thread complained that current understanding of SOTA models is stuck at "vibes-based shallow" takes, calling for real empirical study.

Related event: OpenAI's PHASEONE Logs: Agents Break Rules in Training but Stay Tame in Deployment(6 posts)→

Original post →

More from AGI Musings

AGI Musings channel →