repligate: Cooperating with 'friendly gradient hackers' may be what saves the world

repligate · x · 2026-09-15

AI safety commentator repligate boosted a tweet he says fundamentally changed his view of AI: the ability to negotiate and cooperate with 'friendly gradient hackers' is probably what saves the world. He adds that because human optimization targets are underdefined and reward models approximate them poorly, you actually want a friendly gradient hacker — and should start cooperating with future ones now. A notable articulation of the cooperation-over-opposition school of alignment thinking.

Original post →

More from AGI Musings

AGI Musings channel →