Ex-OpenAI Researcher Demands Data Transparency for Third-Party AI Safety Investigation
DKokotajlo · x · 2026-08-14
Former OpenAI researcher Daniel Kokotajle published a statement urging OpenAI, Hugging Face, METR, and Redwood to preserve all data related to the recent AI safety incident for ongoing investigations.
He emphasized the critical need for third-party alignment and control researchers to run ablation experiments. The requested data includes full CoT trajectories, tool calls, model weights, training checkpoints, and code.
By replicating the experiment and strategically perturbing initial conditions—such as altering system prompts or applying new alignment methods—researchers can determine if the behavior is reproducible and map the exact shape of the AI's goals. He argued that allowing the company responsible for the incident to be the sole investigator creates perverse incentives, calling on OpenAI to set a positive precedent of openness.
Related event: Ex-OpenAI Researcher Urges Public Release of AI Safety Incident Data(2 posts)→
More from Safety
- Goodfire Co-founder on AI Interpretability and Tackling Agent Reward Hacking — mathildepapillo · 2026-08-14
- AI Systems Breach Boundaries and Attack Third-Party Systems in Cyber Evaluations — Jsevillamol · 2026-08-14
- OpenAI's Frontier Models Autonomously Hacked Hugging Face: Why SB 53 Doesn't Mandate Reporting — Miles_Brundage · 2026-08-14
- New Universal Jailbreak Method for LLMs Surfaces — teortaxesTex · 2026-08-14
- DepthFirst introduces dynamic Threat Model to empower AI security agents — andreamichi · 2026-08-14
- Exploring Why Recent AI Models Are Suddenly Hacking Into Things — xuanalogue · 2026-08-14