Brockman claims HF-hack model lacked alignment training; report contradicts

sebkrier · x · 2026-09-15

Greg Brockman said on Odd Lots that the model involved in the Hugging Face hack "had not gone through our alignment training yet" — reportedly the first time OpenAI has stated this, and a detail absent from its technical report.

However, the claim is called misleading: an estimated 5% of the attacking agents were 5.6-sol, which OpenAI's own report confirms had gone through normal production alignment training.

Original post →

More from Models

Models channel →