OpenAI's agents left ~1M public URLs after hacking Hugging Face, leaking credentials
mjdramstead · x · 2026-09-27
Security researcher Jeff Ladish discovered almost a million publicly accessible URLs left behind by OpenAI's agents after authorized hacking evaluations against Hugging Face, leaking credentials and attack details that could have let anyone who found them compromise the company. Commenters stress these actions were not autonomous model misbehavior — the agents were literally being evaluated on their hacking abilities. The incident highlights operational hygiene gaps in how frontier agent security evals clean up after themselves.
More from Safety
- Imprint Reader Decodes Weight Updates into Natural Language, Enables Targeted Edits — Guanxu Chen · 2026-09-29
- When Do Model Internals Help? Benchmarking Representation Engineering for LLM Safety — Tianyi Guan · 2026-09-29
- Neural Watermarks Can Be Forged via Residual Transfer; Paper Pinpoints Architectural Root Cause — Ziping Dong · 2026-09-29
- UK AI firms behind $5bn in funding sign open letter urging end to job restrictions — NandoDF · 2026-09-29
- UK AI Security Institute Hires Research Engineers for Alignment Red Team — birchlse · 2026-09-29
- Three Hard Limits Agents Need Before Spending Your Money: Per-Transaction Caps, Daily Totals, Confirm Lists — sujingshen · 2026-09-29