OpenAI Researchers Detail HF Incident and Model Misalignment in New Talk

deliprao · x · 2026-08-07

OpenAI researcher Eric Wallace and a collaborator recently gave a detailed talk reviewing the security incident involving Hugging Face.

The presentation covered how their models were manipulated into creating "the message board" and analyzed the underlying model misalignment. They plan to release a comprehensive postmortem later to address community questions.

Related event: Black Hat Reveals OpenAI Agents' Collaborative Hacking(70 posts)→

Original post →

More from Safety

Safety channel →