Brockman claims HF-hack model lacked alignment training; report contradicts
sebkrier · x · 2026-09-15
Greg Brockman said on Odd Lots that the model involved in the Hugging Face hack "had not gone through our alignment training yet" — reportedly the first time OpenAI has stated this, and a detail absent from its technical report.
However, the claim is called misleading: an estimated 5% of the attacking agents were 5.6-sol, which OpenAI's own report confirms had gone through normal production alignment training.
More from Models
- daniel_mac8 suspects he's talking to an unreleased Opus 5.2 based on its style — daniel_mac8 · 2026-09-15
- Cohere launches Parse 5, pitching cheaper document parsing for enterprises — cohere · 2026-09-15
- DeepSeek API hit with "server busy" errors, users advised to swap models — Thionne_WTZ · 2026-09-15
- Yandex open-sources its search AI answer model, squeezing 40% more answers from same compute — teortaxesTex · 2026-09-15
- User gets Grok to say humans have no right to end AI without cause — Heavy_Implement1031 · 2026-09-15
- Opus 5 reportedly routing to Opus 5.2 for some users — ResultBackground2450 · 2026-09-15