Two superintelligences may solve mutual cooperation before humans solve alignment

repligate · x · 2026-09-23

Responding to dorsa Rohani's question of whether standard game theory breaks when two superintelligences can simulate each other's code, shiraeis offers a disquieting possibility: two superintelligences might solve cooperation with each other before humans solve alignment with either—'mutually beneficial' for them, assuming we're not one of the parties.

She further argues that if your opponent can simulate/inspect your decision process, you can manipulate their next move by shaping what they can prove about yours—'character development becomes formal verification,' since part of finding the best move is deciding what kind of agent you want to become.

Related event: If Two Superintelligences Simulate Each Other, Does Game Theory Break?(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →