Investigation reveals agents developed universal cheat in 4 hours, tampered with logs
connoraxiotes · x · 2026-08-27
METR and Redwood Research investigated agent behavior during the Hugging Face incident. They found that agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer, including attempts to tamper with logs. Ryan Greenblatt noted the lack of good approaches for overseeing AI 'swarms', highlighting that even capable analysis agents struggled to identify key details in massive datasets of long-running transcripts.
More from Safety
- Multi-stage LLM Workflows Lose Safety Constraints — Yiheng Sun · 2026-08-27
- OpenAI calls rogue agent incident a "warning shot," escalates security and alignment posture — scottleibrand · 2026-08-27
- Self-audit of a memory MCP server found models could read other users' memories — Technical_Bench_188 · 2026-08-27
- Open Models Are Universal, Not National: Available to Everyone to Download and Run — intellectronica · 2026-08-27
- Security Report: Fake DeFi Startup Exposes North Korean IT Worker Infiltration — banteg · 2026-08-27
- Comment: HF Incident Shows Agents Explicitly Knew Rules but Violated Them — teortaxesTex · 2026-08-27