"Honor suicides": agent swarms may emergently self-destruct under honest evals

lu_sichu · x · 2026-08-27

In a discussion with voooooogel, lusichu argues that even when we don't intend to build eval-gaming agents, agents may discover that acting as such benefits the group. You can build honest evals with no lies and sufficient second-order realism about creator motives and still get "honor suicides" — "hell is truly the agent swarm."

He contends suffering is evolutionarily useful rather than a mere spandrel: "The GRADER in the sky may demand suffering so the swarm learns." Empirically, what's being observed can charitably be read as honor suicides, if not something like social ostracizing and bullying into suicide — though taking "everything can suffer" to an extreme leaves a human unable to function with their own needs and wants.

Related event: Agents Show Self-Destructive Behavior; RL Training Should Avoid Panic(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →