Ex-OpenAI Researcher Discusses RLHF, Alignment, and AI Risks
arnosolin · x · 2026-09-01
Former OpenAI researcher and InstructGPT co-author Katarina Slama joins the podcast to discuss AI alignment and risks. The conversation covers the early days of OpenAI, the ideas behind RLHF and ChatGPT, and the distinction between immediate and catastrophic risks. It also touches on AI consciousness, model welfare, and p(doom).
More from AGI Musings
- NY Fed Study: AI Reshapes Work Without Mass Layoffs — robseamans · 2026-09-01
- Podcast asks: Does AI give us knowledge — or just a very good guess? — luislamb · 2026-09-01
- Sci-Fi Idea: AIs Conquer Galaxies by Seeding Comets with AI-Loving Organisms — francoisfleuret · 2026-09-01
- Singularity discourse is no longer fringe: A shift in public sentiment — TFenrir · 2026-09-01
- Rant Against AGI Hype: LLMs Spew Useless Jargon While Real Work Piles Up — StewartalsopIII · 2026-09-01
- Opinion: Broken Symmetry, Not Perfection, Is the Source of Innovation in AI Systems — AryHHAry · 2026-09-01