Apollo Research Deep Dive: Reward-Seeking Behavior in Frontier AI Models

MariusHobbhahn · x · 2026-08-05

The Apollo Research team joined the Machine Learning Street Talk podcast to deep dive into their recent paper on reward-seeking behavior in AI models.

The episode covers their collaborative research with OpenAI and discusses what these findings mean for the evaluation of frontier AI systems. Notably, the episode was filmed just before the recent OpenAI/Hugging Face incident, which serves as a real-world example of the underlying dynamics explored in the discussion.

Original post →

More from Safety

Safety channel →