Apollo Research Explains Testing AI for Hidden Goals: Contrastive Belief Updates

Machine Learning Street Talk · youtube · 2026-08-01

In this MLST episode, Apollo Research discusses their paper on measuring reward-seeking via contrastive belief updates, exploring whether AI can do the right thing for the wrong reason, and how to detect scheming behavior.

Original post →

More from Safety

Safety channel →