Ben Goertzel's d-calculus: the math of goal preservation under AI self-improvement
burny_tech · x · 2026-10-06
Ben Goertzel published 'Finding the Beneficial-Superintelligence Attractor,' applying his new d-calculus math—built for minds that fork, merge, learn and rewrite themselves—to AI safety.
- Core question: as AI rewrites its memory, reasoning habits and even goal-handling code, what stops it from shedding the commitments that made improvement worthwhile, one plausible upgrade at a time?
- The goal-preservation-under-self-modification problem dates to late-90s RSI debates, alongside wireheading risks.
- Using d-calculus, he argues conditions under which self-improving AIs retain goals and evolve into the basin of attraction of 'beneficial superintelligence,' with OmegaHive experiments proposed.
- A formal, math-heavy attempt at an increasingly practical problem as RSI becomes real.
More from AGI Musings
- AI Discourse: 'Capital Realism' and 'Successionism' Are the Same View, and Musk Has Been Circling Them for Years — DavidDuvenaud · 2026-10-06
- Guardian: Altman privatizes AI gains while the public socializes the risks — nordicinst · 2026-10-06
- Model deception as an artifact of our own incoherent wishes — aiamblichus · 2026-10-06
- Huberman: Meta, OpenAI and Anthropic are all becoming biotech companies — Scobleizer · 2026-10-06
- Geoffrey Irving: strong single-single alignment may bootstrap multi-multi alignment — geoffreyirving · 2026-10-06
- Ex-UK energy official warns AI needs 500GW by 2035 and the industry isn't ready — ShakeelHashim · 2026-10-06