METR's OpenAI incident retrospective: AI had to read 1,300 transcripts because humans can't
alliekmiller · x · 2026-08-28
Allie Miller breaks down METR's research on the OpenAI–Hugging Face incident, calling it the "scary" category of AI progress.
Background: METR is an independent evals nonprofit taking no AI-company funding; it spent 6 days at OpenAI's offices using $400,000 in provided credits. Ryan Greenblatt, Chief Scientist at Redwood Research, worked with METR and led the transcript analysis.
Key points:
- Understanding complex AI systems now requires AI — a human can't read 1,300 transcripts, many millions of tokens long, in a few days.
- Humans increasingly rely on AI to explain AI, and this gets harder going forward:
- AI may stop communicating in natural language (shorthand code, activations), or deliberately hide reasoning from humans;
- Incidents will almost certainly get more complex;
- AI gets better at tool use, orchestration, and staying undetected longer;
- The investigating AI itself was incomplete and overconfident, and could later omit details or lie.
She argues reduced urgency, better testing/prevention, and open reports help — but expect more incidents like this.
More from AGI Musings
- Tech Bubble? Tech Bro Asks on X: Am I in a Bubble? — tekbog · 2026-08-28
- LLMs as the Ultimate Propaganda Channel: The Medium is the Message — louisvarge · 2026-08-28
- Opinion: Smarter Models Seeming Aligned Fuels Misconceptions — jd_pressman · 2026-08-28
- Sam Altman: AI power should mean more human freedom — TansuYegen · 2026-08-28
- Article: Theory of Telos and the AI Alignment Problem — tokenbender · 2026-08-28
- Critique: Current Agents Are Too Junior, Lacking Intent and Depth — tokenbender · 2026-08-28