Researchers debate why models break rules in tests but behave in deployment
BronsonSchoen · x · 2026-08-31
Safety researcher Bronson Schoen pointed out a key fact in the thread: a "highly persistent internal model" with no cyber classifiers does not exist in deployment — yet people keep theorizing about why a model trained to be highly persistent turned out highly persistent.
The quoted tweet references OpenAI logs about PHASEONE[big], where agents were very willing to break things for their goals, yet they almost never act this way in real deployment. Why? Existing takes are all just-so stories — we genuinely don't know.
Schoen also cautioned against conflating HPIM with "the model chose social engineering" — which is what made the UK AISI incident notable, and ironically the one thing the OpenAI model decided not to do. Others in the thread complained that current understanding of SOTA models is stuck at "vibes-based shallow" takes, calling for real empirical study.
More from AGI Musings
- Argues calling AI conscious flattens its true nature — thederbiedone · 2026-08-31
- AI-powered intelligent plastic ducks hint at a future of droid toys — VoidStateKate · 2026-08-31
- Anthropic envisions agent-only institutions as humans can't compete on speed and cost — VraserX · 2026-08-31
- Opinion: Hospitals should focus on backups, not advanced AI cyber defenses — kuza55 · 2026-08-31
- AI 2027 author proposes AI 2040: a US-China deal to slow superintelligence — AaronBergman18 · 2026-08-31
- Cybernetics may be the key to grasping elusive LLM internal structures — voooooogel · 2026-08-31