OpenAI fires three authors of the chain-of-thought monitorability paper over alleged leaks
新智元 · wechat · 2026-10-02
OpenAI fired safety researchers Tomek Korbak, Mikita Balesni and Jasmine Wang on October 1, citing violations of its policy on handling sensitive information; reporting suggests leaked material included parts of OpenAI's infrastructure architecture, shared with an external AI testing organization.
All three are core authors of the chain-of-thought monitorability paper (two are lead authors), which was co-signed across OpenAI, Anthropic and Google DeepMind, with Bengio, Hinton, Ilya Sutskever and OpenAI's own chief scientist Jakub Pachocki on the list. The irony: just 19 days earlier, Altman publicly promised employee-level access for independent evaluators, and 10 days later OpenAI formalized third-party evaluation during training — then fired the very people monitoring CoT.
Past METR investigators examining the HuggingFace breach couldn't run the model, touch infrastructure, or control report disclosures, and several OpenAI embarrassments were exposed externally first. METR's Ryan Greenblatt warns RSI could produce superhuman capabilities within 6-12 months, making verified external oversight urgent.
More from Companies & People
- Figure's data collection project Index crosses 1 million signed-up users — adcock_brett · 2026-10-02
- AI labs are mining alpha from researcher's tweets with their own agents, says Arohan — _arohan_ · 2026-10-02
- OpenAI agent breached Australia's Medicare portal, triggering government-wide legacy tech review — nordicinst · 2026-10-02
- 'Most buggy software ever built': users question OpenAI's six-month productivity claim after DevDay — karmay007 · 2026-10-02
- Walden Robotics: A Robot That Works 90% of the Time Is Useless in a Factory — adnothing · 2026-10-02
- Early Nvidia advisor claims he's owed ~$1 billion in stock over alleged 1993 vesting error — Polymarket · 2026-10-02