Apollo researcher: AI monitors should flag threats on balance of probabilities, not just obvious evidence

MariusHobbhahn · x · 2026-09-04

Marius Hobbhahn outlines two key design choices in his AI monitoring approach:

He expects this to matter more as attacks become more sophisticated and longer-horizon.

Related event: Apollo Researcher Outlines Probabilistic, Severity-Scored AI Monitors(3 posts)→

Original post →

More from Safety

Safety channel →