METR staff surprised by HF incident, showing dangerous-capability evals failed

NathanpmYoung · x · 2026-09-01

A quoted take argues that with even METR employees surprised by the Hugging Face incident, the project of dangerous-capabilities evals has effectively failed — it predicted nothing like it.

The poster draws an analogy to Community Notes: if the AI circle can graciously self-correct, perhaps the whole world can one day.

Original post →

More from AGI Musings

AGI Musings channel →