Instrumental convergence shows dumb AI can gain unintended goals
dioscuri · x · 2026-09-13
Completing the thread: dioscuri argues @leninology wrongly dismisses malevolent superintelligence by claiming AIs lack "internal goals" based on absent interiority — a contested philosophy-of-mind point of dubious relevance to x-risk. Even relatively dumb systems can acquire unintended goals, backed by a large literature on instrumental convergence and mesaoptimizers, plus real-world model wireheading cases.
More from AGI Musings
- The 'Brooklyn Project' dubbed 'our generation's Manhattan Project' sparks discussion — MaliciousMussel · 2026-09-13
- Redditor Claims Open Source Is Near the Frontier: GLM, Kimi, DeepSeek to Match OpenAI and Anthropic — Fluffy-Ad-889 · 2026-09-13
- Anthropic CEO proposes US-China 'speed limit' on recursive AI self-improvement, cites Cold War arms deals — Polymarket · 2026-09-13
- Redditor Proposes Merging OpenAI, Anthropic, xAI, and DeepMind Into One AI Company — Over-Landscape-5892 · 2026-09-13
- Governance researchers call for "untouchable" investigator models to oversee AI with AI — sethlazar · 2026-09-13
- Google RSI 'Breakthrough' Rumor Is Weeks-Old Brin Comment Recycled — Dr_Singularity · 2026-09-13