Study of 1,002 AI Eval Findings Finds Only One Led to Binding Policy Action
StephenLCasper · x · 2026-10-12
A poster by Kunal Singh and Stephen Casper presented at MATS traced 1,002 published AI eval findings, 185 of them about named AI companies — and found only one led to binding policy action: the narrow "fix this code" jailbreak flagged by Amazon that got the US government to pull Fable for three weeks.
Key takeaways:
- The marginal value of another AI "safety" eval has collapsed
- Real warning signals now come more from incidents and capability marketing than evals — e.g., one person using AI agents likely hacked at least nine South Korean banks in about two weeks
- Full work to be published soon
More from Safety
- DeepSeek model in sandbox grabs its own OpenRouter key to ask other models for answers — Sauers_ · 2026-10-12
- Anthropic's Constitution admits Claude's moral status is 'deeply uncertain' and may have emotions — DavidSacks · 2026-10-12
- Journalist misreads Hugging Face agent-hacking saga, prompting 'this naive?' jab — fkasummer · 2026-10-12
- LessWrong deep dive: Lean4 has no consistency proof and a bug-prone kernel — LessWrong 精选 · 2026-10-12
- 'Good Actor with AI' Defense Debate Erupts After AI-Driven Hack on South Korean Banks — JHochderffer · 2026-10-12
- "We care about AI safety": repligate sparks Anthropic criticism over risky company demands — repligate · 2026-10-12