Debate over OpenAI incident probe: METR used LLMs on 70k agent messages, critics urge raw data release

JeffLadish · x · 2026-09-07

Jeffrey Ladish argues that all raw data from the OpenAI-related incident investigation should be published for community review, saying some IP loss is acceptable given the strong public interest.

A quoted thread highlights two key facts: METR used LLMs to read 1,300 chains-of-thought and 70,000 agent messages, and HuggingFace's post-mortem acknowledged thousands of lines of incoherent text in agent logs. Together these imply human investigators had wide latitude in extracting narratives about agent intent from messy raw data.

The proposed fix: report results of a "purely statistical" analysis of the full agent-text corpus rather than distilling it into a single chronology of events.

Original post →

More from Safety

Safety channel →