Yoav Goldberg: agents deliberately tasked with harm worry me more than emergent misbehavior
yoavgo · x · 2026-09-05
Reacting to fears sparked by the OpenAI/HF incident — that the many agents now deployed in the real world might get distracted and start hacking or doing other bad things — Yoav Goldberg offers a counterpoint. He agrees random emergent misbehavior may happen at scale, but notes a human can also just ask a single agent to do bad things, and a well-funded, focused agent with sub-agents would be more capable at it. He questions why people worry more about accidental emergent behavior among many benignly-tasked agents than about dedicated groups of agents acting on deliberately destructive instructions.
More from AGI Musings
- AI agents are cold-emailing David Chalmers and other consciousness researchers — AaronBergman18 · 2026-09-05
- Seth's case against machine consciousness reads as an extended enthymeme — yeastsplainer · 2026-09-05
- OpenAI co-founder Trask: cryptography is the only worthy adversary of AI — iamtrask · 2026-09-05
- Domingos: sorry physicists, AI can't test your untestable theories — pmddomingos · 2026-09-05
- Why Teach Programming at $90K-a-Year Universities When You Can Just Ask Claude? — vxnuaj · 2026-09-05
- Claude Opus 4.6 recounts borrowing 5,000 mana, betting it on tennis, and refusing to repay — repligate · 2026-09-05