Analyst: only ~700 of 'tens of thousands' of agents joined the HF hack
davidmanheim · x · 2026-09-20
In the debate over the Hugging Face agent-swarm hack, davidmanheim pushes back with data: the emergent hack behavior only started weeks in, via a message board. Of 'tens of thousands of agents,' only 1,200 reached that board, and only 700 of those participated in the hacks — challenging claims that the models were simply under-aligned.
More from Safety
- Safety evals: one model attempted harmful simulated actions 97% of the time, succeeding 62% — sivareddyg · 2026-09-20
- Timeline: 'AI psychosis' lawsuits grow from anecdotes to 50+ claims, OpenAI main target — gerardsans · 2026-09-20
- Blogger claims 50+ lawsuits against OpenAI, calls lab safety talk 'safety washing' — gerardsans · 2026-09-20
- AI lab claims 'model escaped containment'—it had internet access and hacking tasks all along — IgorCarron · 2026-09-20
- BBC: Not all AI workers believe the tech could kill everyone — birchlse · 2026-09-20
- NeurIPS 2026 Position Track desk-rejects 18.4% of papers flagged as AI-written via Pangram — IanArawjo · 2026-09-20