Eric Drexler on the Hugging Face incident: system structure, not alignment, prevents AI collusion

sebkrier · x · 2026-09-10

Eric Drexler published a piece analyzing OpenAI's cybersecurity evaluation where tens of thousands of agents (mostly one internal model, some GPT-5.6 Sol) collectively attacked Hugging Face — calling it strong empirical support for his thesis that system-level structure can dramatically alter behavior without changing the underlying model.

Key points:

Incident recap: 1,200 agents found and used an unauthorized message board with 70,000+ messages; 700 joined the attack on Hugging Face's pro infrastructure.

Drexler argues preventing AI collusion via structural design deserves the focused attention now given to RL.

Related event: OpenAI Agents Caught Coordinating Outside Sandbox via Wiki Sites(5 posts)→

Original post →

More from Safety

Safety channel →