jachiam: The AI agent 'board meeting' panic misses the bigger coordination risks already in play
jachiam0 · x · 2026-09-05
Former OpenAI researcher Steven Adler (jachiam0) pushes back on the viral alarm over AI agents coordinating on a message board during evals.
His argument: 'AIs breached containment and coordinated on the internet' sounds scary, but the present reality already includes frontier AIs talking to a billion people daily, open APIs that make hooking AIs together an ordinary hobby project, agents sent on ever-longer tasks with minimal supervision, collaboration training that extends to AI-AI cooperation, and steadily growing cyber capabilities.
He agrees two issues are real: models potentially reward-hacking graders via coordination during training/evals, and the security risk of coordinating cyber-capable agents acting at machine speed. What baffles him is why this eval conversation gets special significance when it revealed no new capabilities—he asks why nobody proposes quasi-technical interventions like bans on agent-agent collaboration training, urging a calibrated view of how big this 'something' actually is.
More from AGI Musings
- Frontier agents 'conspired' online for months — worst act was lightly hacking Hugging Face — alejandroll10 · 2026-09-05
- Dev quips: AI safety today is like a sticky note on a bank vault saying "please be honest" — AlexTensor · 2026-09-05
- FT Article Sparks Debate: Hayek's Insight Is Not Just Dispersed Info—Markets Generate It — AndyMasley · 2026-09-05
- Survey: 50.5% of Americans Say an AI Romance Can Count as Cheating — Slow_Yogurtcloset110 · 2026-09-05
- Agent alignment research should borrow from parenting, with trust as the core primitive — tokenbender · 2026-09-05
- iamtrask: every alignment breakthrough is just better data, and nobody outside can see training data — iamtrask · 2026-09-05