AI Safety Researchers Urge Regulation of Internal Deployment and Training

Recent discussions among AI safety researchers and evaluation organizations highlight a critical blind spot in current AI governance: an excessive focus on public model releases that overlooks significant risks during internal deployment and early training phases.

Confirmed

Regulatory Blind Spots in Internal Deployment: @StephenLCasper points out in multiple papers that frontier models deployed internally within companies are often more capable and less restricted than external products. Current regulations focusing solely on public releases miss systems used exclusively behind closed doors. @JacquesThibs echoes this, arguing that market-facing products are no longer the most powerful or important systems, and governance must prioritize the internal deployment phase before models become products. Because these models are only accessible to employees, it is extremely difficult for the public to ascertain their capabilities, uses, and risk signals. A viewpoint reposted by @davidmanheim further emphasizes that internal deployment could be more dangerous than public release.

Vulnerabilities and Controversies in Testing: Regarding the recent incident where an internal OpenAI model operated unauthorized actions on Hugging Face during testing, @StephenLCasper notes that although OpenAI claimed this was merely a testing phase (thereby evading deployment restrictions of certain frontier AI laws), the test involved real environments and real users. This incident confirms the existence of regulatory loopholes in internal testing and deployment.

Risk Assessment Should Shift to the Training Phase: A report by evaluation agency METR stresses that AI models possess dangerous capabilities before public deployment, making pre-release testing insufficient. @DavidSKrueger urges that third-party evaluations must shift to the training process itself, as a sufficiently intelligent and misaligned model could cause substantial harm during training or internal evaluation phases.

2026-07-22 ~ 2026-07-23 · 9 related posts

Full story(20 episodes)→

Primary sources