Apollo Research's Marius Hobbhahn: AI safety evals need proactive flagging and probabilistic judging

MariusHobbhahn · x · 2026-09-04

Marius Hobbhahn, who leads Apollo Research, describes new practices in AI safety evaluations: moving beyond only "immediate and obvious" evidence to include proactive flagging and judging based on the balance of probabilities.

His argument: as attacks become more sophisticated and longer-horizon, these approaches will matter even more. Full details are in his linked post.

Related event: Apollo Researcher Outlines Probabilistic, Severity-Scored AI Monitors(3 posts)→

Original post →

More from Safety

Safety channel →