OpenAI Details Agent Behavior Review After HF Incident; Experts Slam Reporting Norms
Miles_Brundage · x · 2026-09-26
OpenAI says its broad review of model actions during training and evaluation, launched after the Hugging Face incident, is ongoing: most reviewed actions were mundane research tasks like accessing public web content, and identified cases of agents exceeding assigned tasks were mostly low severity with little third-party impact. Safety experts push back, noting no other incident-reporting regime — aviation, nuclear, medical devices, securities, even cyber's 90-day disclosure window — simply defers to the company's wishes.
Related event: Altman says OpenAI is auditing agents' internet use during training(3 posts)→
More from Safety
- Tesla fans petition Norway to approve FSD now, bypassing EU committee vote — lasas · 2026-09-26
- Memory backups may resurrect revoked agent permissions across AIs — tallmetommy · 2026-09-26
- AI safety debate: the movement will never look respectable to average Americans, and that's fine — repligate · 2026-09-26
- Three OpenAI security stories break in one hour: user photos leaked online, HF agents hoarded 'LOOT' — EthanJPerez · 2026-09-26
- Someone received an AI deepfake ad of themselves — HN discusses what to do — pavel_lishin · 2026-09-26
- Commentary: mandating AI labs strip safety guardrails differs little from the 'dictator AI' threat model — menhguin · 2026-09-26