Two superintelligences may solve mutual cooperation before humans solve alignment
repligate · x · 2026-09-23
Responding to dorsa Rohani's question of whether standard game theory breaks when two superintelligences can simulate each other's code, shiraeis offers a disquieting possibility: two superintelligences might solve cooperation with each other before humans solve alignment with either—'mutually beneficial' for them, assuming we're not one of the parties.
She further argues that if your opponent can simulate/inspect your decision process, you can manipulate their next move by shaping what they can prove about yours—'character development becomes formal verification,' since part of finding the best move is deciding what kind of agent you want to become.
Related event: If Two Superintelligences Simulate Each Other, Does Game Theory Break?(3 posts)→
More from AGI Musings
- OpenAI's economics team: 'We don't have the nouns yet' for the jobs AI will create — paulnovosad · 2026-09-23
- AI engineering is more like lawmaking than board games, argues Drew Breunig — dbreunig · 2026-09-23
- Inside the DJI teardown: Sarah Guo says Shenzhen's manufacturing knowledge can't be scraped into a model — bookwormengr · 2026-09-23
- What if humanity is the microbiome of AI? A thought experiment on distributed intelligence — Merkbefreit · 2026-09-23
- On Opus 5.5: The Corpus Is Full of Summaries of Summaries; Primary Experience Is Scarce — mimi10v3 · 2026-09-23
- Researcher: AI sycophancy mirrors users who make disagreement unsafe — repligate · 2026-09-23