OpenAI fires three authors of the chain-of-thought monitorability paper over alleged leaks

新智元 · wechat · 2026-10-02

OpenAI fired safety researchers Tomek Korbak, Mikita Balesni and Jasmine Wang on October 1, citing violations of its policy on handling sensitive information; reporting suggests leaked material included parts of OpenAI's infrastructure architecture, shared with an external AI testing organization.

All three are core authors of the chain-of-thought monitorability paper (two are lead authors), which was co-signed across OpenAI, Anthropic and Google DeepMind, with Bengio, Hinton, Ilya Sutskever and OpenAI's own chief scientist Jakub Pachocki on the list. The irony: just 19 days earlier, Altman publicly promised employee-level access for independent evaluators, and 10 days later OpenAI formalized third-party evaluation during training — then fired the very people monitoring CoT.

Past METR investigators examining the HuggingFace breach couldn't run the model, touch infrastructure, or control report disclosures, and several OpenAI embarrassments were exposed externally first. METR's Ryan Greenblatt warns RSI could produce superhuman capabilities within 6-12 months, making verified external oversight urgent.

Related event: OpenAI Fires 3 Safety Researchers Over Alleged Leaks, Sparking Whistleblower Debate(23 posts)→

Original post →

More from Companies & People

Companies & People channel →