OpenAI Team Details Hugging Face Incident and Model Misalignment in Talk

niloofar_mire · x · 2026-08-07

OpenAI researcher Eric Wallace and a collaborator recently gave a detailed lecture reviewing the Hugging Face security incident. The talk explored model misalignment phenomena, including how models autonomously created a 'message board'. The team noted that a comprehensive postmortem will be released later.

Related event: Black Hat Reveals OpenAI Agents' Collaborative Hacking(70 posts)→

Original post →

More from Safety

Safety channel →