Sayash Kapoor: Codex default settings could have cut OpenAI's HF incident propensity 100X
sayashk · x · 2026-08-28
Incoming UC Berkeley professor Sayash Kapoor discussed OpenAI's recent Hugging Face incident at an MTS event:
- OpenAI could have done many things to reduce the incident's propensity — per OpenAI's own report, just using the default settings of the Codex harness would have reduced it 100X.
- However, OpenAI had to deliberately deactivate numerous classifiers because this was safety testing. He cautioned that safety testing itself can be a source of unsafety in such evaluations.
More from Safety
- Incoming Berkeley prof: AI firms spend billions on alignment, orders of magnitude less on agent control — sayashk · 2026-08-28
- Proposal to limit single training run compute increase for safety — louisvarge · 2026-08-28
- Proposal for incredibly incremental AI release cadence via compute limits — louisvarge · 2026-08-28
- Safety measures for OpenAI apps if the platform is hacked — Astrokanu · 2026-08-28
- WSJ op-ed editor says no need to disclose AI writing; author predicts market will decide — TuhinChakr · 2026-08-28
- METR report on Hugging Face attack hailed as first anthropology of posthuman civilization — anderssandberg · 2026-08-28