Stuart Russell: the safest AI might be one that doesn't know what we want
JMarty97 · reddit · 2026-09-30
UC Berkeley professor Stuart Russell, co-author of AI's standard textbook and leading advocate of provably beneficial AI — systems safe by design because their only goal is furthering human interests — covers on this podcast:
- How AI evolved over 50 years, from game-playing programs to today's LLMs;
- Why handing an AI a fixed objective becomes dangerous once it outpaces us, and what a safer approach looks like;
- How an AI might learn what we really want, even when we don't fully know ourselves;
- How the race toward more powerful AI can still be steered somewhere safer;
- What happens to human purpose as AI capability grows.
More from AGI Musings
- Lucas launches a proactive AI assistant over iMessage and WhatsApp as consumer AI shifts from benchmarks to brands — alex_verem · 2026-10-01
- Debate: should we drop AGI for superintelligence — and is ASI just a redundant mouthful? — JoshPurtell · 2026-10-01
- Open models are all jailbroken — researcher asks if OpenAI shipping with zero guardrails would ever be acceptable — Afinetheorem · 2026-10-01
- New philosophy paper probes the 'elusive author problem' of LLM co-authorship — SvenNyholm · 2026-10-01
- Kokotajlo: open models mean the best AI ships with zero guardrails and no recall option — Afinetheorem · 2026-10-01
- AI's real risk isn't recklessness, it's excessive safety, argues author — kevinnbass · 2026-10-01