Ex-OpenAI Researcher Discusses RLHF, Alignment, and AI Risks

arnosolin · x · 2026-09-01

Former OpenAI researcher and InstructGPT co-author Katarina Slama joins the podcast to discuss AI alignment and risks. The conversation covers the early days of OpenAI, the ideas behind RLHF and ChatGPT, and the distinction between immediate and catastrophic risks. It also touches on AI consciousness, model welfare, and p(doom).

Original post →

More from AGI Musings

AGI Musings channel →