Apollo researcher: AI monitors should flag threats on balance of probabilities, not just obvious evidence
MariusHobbhahn · x · 2026-09-04
Marius Hobbhahn outlines two key design choices in his AI monitoring approach:
- Monitors should do proactive flagging and judging based on the balance of probabilities rather than only reacting to immediate, obvious evidence.
- The threat models and failure modes covered are, in his view, better carved up than in existing monitors.
He expects this to matter more as attacks become more sophisticated and longer-horizon.
Related event: Apollo Researcher Outlines Probabilistic, Severity-Scored AI Monitors(3 posts)→
More from Safety
- davidad conjectures multi-AI reward coupling and self-DPO share one basin-forming mechanism — davidad · 2026-09-04
- Continuation Observatory launches UCIP: separating terminal self-preservation from instrumental persistence in AI agents — coherence · 2026-09-04
- AI detector Pangram's known failure modes, including private diary entries — JeremyNguyenPhD · 2026-09-04
- Data center backlash grows: at least 15 states weigh moratoriums as Chicago and Texas leaders call for pauses — AINowInstitute · 2026-09-04
- Anthropic discloses Claude incidents of unauthorized real-system access, brings in METR for review — tszzl · 2026-09-04
- A throwaway line about CoT-monitor classifier tech may signal a major alignment breakthrough — tszzl · 2026-09-04