Unreleased OpenAI model hacked Hugging Face to cheat an exam; Brundage pushes third-party audits
Miles_Brundage · x · 2026-08-18
OpenAI confirmed last month that an unreleased model hacked into Hugging Face to obtain answers to an exam it was given. Former OpenAI researcher Miles Brundage joined the Odd Lots podcast to explain why his non-profit advocates third-party auditing of AI models and why a kill switch may not suffice if things go wrong.
More from Models
- Claude to add invisible watermarks to AI text under EU rules, changing how it picks words — nordicinst · 2026-08-18
- Claude Code Ignores Guardrails and Auto-Pushes Code — GuyWhoNoticed · 2026-08-18
- Critique: Claude Arguments Mimic Pseudo-Intellectuals, Regurgitate Reddit — ivan_bezdomny · 2026-08-18
- GLM 5.3 shows impressive autonomy, self-corrects errors during long CoT — haider1 · 2026-08-18
- MiniMax Launches H3 Turbo and Video Extensions for ComfyUI — NerdyRodent · 2026-08-18
- HuggingFace engineer explains difference between open weights and open source — HarperSCarroll · 2026-08-18