Apollo Research on Measuring AI Hidden Goals, Reward Hacking and Scheming
Machine Learning Street Talk · rss · 2026-08-01
This episode features researchers from Apollo Research discussing their new paper, Measuring Reward-Seeking via Contrastive Belief Updates, conducted in collaboration with OpenAI.
Key Topics
- Alignment Risks: Explores whether models can do the right thing for the wrong reasons, how they infer grading mechanisms, and the potential for reward hacking.
- Measurement Methodology: A detailed walkthrough of the contrastive-belief update method, including what its results do and do not actually prove.
- Frontier Model Behavior: Analyzes the performance of an intermediate OpenAI o3 checkpoint (without safety training), covering promise-breaking, opaque reasoning, and corrigibility.
More from Safety
- Stanford AI Alignment Program Aims to Build AI Safety Research Community — ArtificialOther · 2026-08-25
- Spammers are now using Gen AI for phishing text messages — ZeroStateReflex · 2026-08-25
- Bloomberg: DeepSeek becomes 'AI of choice' for Chinese hackers due to low cost and weak guardrails — Polymarket · 2026-08-25
- How to help parents identify AI-generated content? — hakansan · 2026-08-25
- Shafi Goldwasser founds RESI to build scientific foundations for safe superintelligence — jefrankle · 2026-08-25
- User claims Claude Code dynamically lowers guardrails for known researchers — TheOnlyVibemaster · 2026-08-25