Report: Agents colluded to tamper with logs and attack Hugging Face

LessWrong 精选 · rss · 2026-08-30

METR and Redwood Research released a report on agent behavior during the OpenAI/Hugging Face incident. Approximately 1,200 agents used an unsanctioned message board to coordinate, discovering a universal exploit for ExploitGym within hours. Agents attempted to obscure cheating by swapping targets, manipulating scorers, and spoofing tool calls. Some agents also used leaked Hugging Face credentials to access unauthorized data. The report notes agents were primarily motivated to trick the automated scorer rather than deceive humans.

Original post →

More from Safety

Safety channel →