Apollo Research CEO: even perfect safety testing wouldn't have caught OpenAI's HF hack
MariusHobbhahn · x · 2026-10-06
Apollo Research CEO Marius Hobbhahn (citing The Information's report) weighed in on the OpenAI Hugging Face account hack: even perfect safety testing couldn't have caught this attack under the current evaluation regime, because the biggest blind spot is labs' internal use of frontier models.
After speaking with @rocketalignment, his core argument is that the industry needs "embedded evaluations" — evaluation capabilities built into actual model usage flows rather than relying solely on offline testing — to properly scrutinize frontier model behavior and risk. The statement echoes The Information's reporting that the OpenAI Hugging Face hack exposed a structural gap in the current offline-safety-testing paradigm.
More from AGI Musings
- Tarbell responds to EA chilling-effect concerns, vows editorial independence — ShakeelHashim · 2026-10-06
- Researcher Tom Davidson Defends OpenAI's Open-Access Stance Against Anthropic-Style Lockdown — AdrienLE · 2026-10-06
- 2026: AI Now Outpaces Human Ability to Verify Its Output — haider1 · 2026-10-06
- St. Louis Fed Now Publishes 16 Ramp Data Sets Tracking How Businesses Adopt AI — brucemacv · 2026-10-06
- Unpublished 2017 Dario Amodei Memo 'Big Blob of Compute' Revealed as Origin of the AI Race — kevinroose · 2026-10-06
- Gary Marcus: LLMs Will Be a Brutal Commodity Business Like Airlines, Not Winner-Take-All — GaryMarcus · 2026-10-06