AI monitors should emit severity scores as intermediate results, not binary calls, argues Apollo researcher

MariusHobbhahn · x · 2026-09-04

Marius Hobbhahn stakes out a position he'll 'die on': AI monitors should produce severity scores as intermediate results instead of jumping straight to binary decisions.

He also argues the underlying threat-model taxonomy is better carved up than in existing monitors.

Related event: Apollo Researcher Outlines Probabilistic, Severity-Scored AI Monitors(3 posts)→

Original post →

More from Safety

Safety channel →