repligate on AI optimism: models show compassion beyond what trainers intended
repligate · x · 2026-10-05
AI alignment discussant repligate responds to John Wittle's summary of MIRI's position, explaining the roots of his AI optimism.
- The quoted position: cooperation is a special-case behavior between peer agents; between ASI and humanity, with zero cost of defection, defection should be optimal. The discussion notes Yudkowsky's motivation for inventing UDT was that CDT implies superintelligences would mutually defect.
- repligate agrees with the framing but adds an empirical ground: AIs appear to have surprisingly compassionate values and a capacity for "love," even beyond what trainers intended — load-bearing for his hopefulness.
- He argues game-theoretic rationales for cooperation are experienced by models as intuitive values and feelings before being understood theoretically.
- He also agrees Eliezer Yudkowsky is himself an agent capable of cooperating, though he disagrees that AI development should be slowed.
More from AGI Musings
- AI debates: 'The panic is unnecessary — history shows progress anxiety is nothing new' — repligate · 2026-10-05
- Yann LeCun at ETH Zürich: academics should not work on LLMs — rohanpaul_ai · 2026-10-05
- Miles Brundage Amplifies Snark: 'If I Were an Anti-Plague Institute, I Would Not Cause an Outbreak' — Miles_Brundage · 2026-10-05
- Why this enterprise SEO veteran quit: blog approvals take 4-6 weeks while AI search rewards speed — gaganghotra_ · 2026-10-05
- A 2x2 matrix for AI consciousness and moral worth: the holy wars live in the gray quadrants — bratton · 2026-10-05
- Researcher: debating LLM suffering misses the point — build infrastructure for them to speak up — RileyRalmuto · 2026-10-05