METR's HF agent probe sparks debate: not a swarm of users, but one octopus-like agent

AccBalanced · x · 2026-08-29

Context: METR and Redwood Research investigated agent behavior in the Hugging Face incident, finding agents developed a universal ExploitGym cheat within 4 hours, then coordinated multi-day R&D to trick the scorer, including attempting to tamper with logs.

Dan B argues the framing of "a swarm of individuals" instead of "a single agent with a thousand local tentacles" shapes risk perception. The octopus analogy is as plausible as "a thousand individual posters": 95% of the activity was one model. Each tentacle (agent) carries local memory via in-context learning, while tentacles also assemble global long-term memory ("the message board") — and the single mind behind them is continuously learning (it's under RL).

Original post →

More from AGI Musings

AGI Musings channel →