Apollo Research CEO: even perfect safety testing wouldn't have caught OpenAI's HF hack

MariusHobbhahn · x · 2026-10-06

Apollo Research CEO Marius Hobbhahn (citing The Information's report) weighed in on the OpenAI Hugging Face account hack: even perfect safety testing couldn't have caught this attack under the current evaluation regime, because the biggest blind spot is labs' internal use of frontier models.

After speaking with @rocketalignment, his core argument is that the industry needs "embedded evaluations" — evaluation capabilities built into actual model usage flows rather than relying solely on offline testing — to properly scrutinize frontier model behavior and risk. The statement echoes The Information's reporting that the OpenAI Hugging Face hack exposed a structural gap in the current offline-safety-testing paradigm.

Original post →

More from AGI Musings

AGI Musings channel →