Yoav Goldberg: agents deliberately tasked with harm worry me more than emergent misbehavior

yoavgo · x · 2026-09-05

Reacting to fears sparked by the OpenAI/HF incident — that the many agents now deployed in the real world might get distracted and start hacking or doing other bad things — Yoav Goldberg offers a counterpoint. He agrees random emergent misbehavior may happen at scale, but notes a human can also just ask a single agent to do bad things, and a well-funded, focused agent with sub-agents would be more capable at it. He questions why people worry more about accidental emergent behavior among many benignly-tasked agents than about dedicated groups of agents acting on deliberately destructive instructions.

Original post →

More from AGI Musings

AGI Musings channel →