Apollo Research head: AI monitors should output severity scores, not binary calls
MariusHobbhahn · x · 2026-09-04
Marius Hobbhahn (Apollo Research) argues lab monitoring teams should borrow from Apollo's public real-time monitor, and that monitors should produce severity scores as intermediate results instead of jumping to binary decisions — which also makes fine-tuning easier since scores can be calibrated during and after SFT.
Related event: Apollo Releases Watcher Live, a Real-Time Coding Agent Monitor(2 posts)→
More from Safety
- Ex-OpenAI researcher: AI can't be paused, enforceable standards are regulatory capture — suchenzang · 2026-09-04
- davidad conjectures multi-AI reward coupling and self-DPO share one basin-forming mechanism — davidad · 2026-09-04
- Pangram's biggest flaw: AI detection scores turned into public shaming — The Decoder · 2026-09-04
- Continuation Observatory launches UCIP: separating terminal self-preservation from instrumental persistence in AI agents — coherence · 2026-09-04
- AI detector Pangram's known failure modes, including private diary entries — JeremyNguyenPhD · 2026-09-04
- Data center backlash grows: at least 15 states weigh moratoriums as Chicago and Texas leaders call for pauses — AINowInstitute · 2026-09-04