Apollo Research Explains Testing AI for Hidden Goals: Contrastive Belief Updates
Machine Learning Street Talk · youtube · 2026-08-01
In this MLST episode, Apollo Research discusses their paper on measuring reward-seeking via contrastive belief updates, exploring whether AI can do the right thing for the wrong reason, and how to detect scheming behavior.
More from Safety
- Stanford AI Alignment Program Aims to Build AI Safety Research Community — ArtificialOther · 2026-08-25
- Spammers are now using Gen AI for phishing text messages — ZeroStateReflex · 2026-08-25
- Bloomberg: DeepSeek becomes 'AI of choice' for Chinese hackers due to low cost and weak guardrails — Polymarket · 2026-08-25
- How to help parents identify AI-generated content? — hakansan · 2026-08-25
- Shafi Goldwasser founds RESI to build scientific foundations for safe superintelligence — jefrankle · 2026-08-25
- User claims Claude Code dynamically lowers guardrails for known researchers — TheOnlyVibemaster · 2026-08-25