Debate over OpenAI incident probe: METR used LLMs on 70k agent messages, critics urge raw data release
JeffLadish · x · 2026-09-07
Jeffrey Ladish argues that all raw data from the OpenAI-related incident investigation should be published for community review, saying some IP loss is acceptable given the strong public interest.
A quoted thread highlights two key facts: METR used LLMs to read 1,300 chains-of-thought and 70,000 agent messages, and HuggingFace's post-mortem acknowledged thousands of lines of incoherent text in agent logs. Together these imply human investigators had wide latitude in extracting narratives about agent intent from messy raw data.
The proposed fix: report results of a "purely statistical" analysis of the full agent-text corpus rather than distilling it into a single chronology of events.
More from Safety
- LLM vulnerability-fixing pipelines introduce bugs more than they fix, study says — dyn___ · 2026-09-07
- Should humans intervene in an alien civilization's path to its own singularity? — jachiam0 · 2026-09-07
- Podcast on AI safety digs into the recent OpenAI / Hugging Face attack — arnosolin · 2026-09-07
- Australia to require social apps to let users turn off algorithms and see only followed posts — santoshpanda · 2026-09-07
- Repeat-After-Me: Black-Box Visual Prompt Injection Hits 47% ASR on GPT-5.5, 80%+ on Open VLMs — chaumian · 2026-09-07
- doodlestein responds to Jakub's alignment essay with 'FrankenAlignment' agent toolbox plan — doodlestein · 2026-09-07