AI monitors should emit severity scores as intermediate results, not binary calls, argues Apollo researcher
MariusHobbhahn · x · 2026-09-04
Marius Hobbhahn stakes out a position he'll 'die on': AI monitors should produce severity scores as intermediate results instead of jumping straight to binary decisions.
- Intermediate scores allow layered review and downstream triage.
- They make fine-tuning monitors easier: you can calibrate both during and after SFT.
He also argues the underlying threat-model taxonomy is better carved up than in existing monitors.
Related event: Apollo Researcher Outlines Probabilistic, Severity-Scored AI Monitors(3 posts)→
More from Safety
- OpenAI rolls out GPT-6 Astra to vetted cyber customers at 2.5x GPT-5.6 pricing — rohanpaul_ai · 2026-09-04
- Podcast breaks down METR and OpenAI reports on the Hugging Face 'swarm' — Gregory_C_Allen · 2026-09-04
- OpenAI urges shared AI safety standards, pressed on why it isn't leading them — RebeccaBellan · 2026-09-04
- TheZvi's AI #184: Five HuggingFace Hack Postmortems and the New Most Capable Model — TheZvi · 2026-09-04
- Virginia State Study: Most Data Centers Use No More Water Than a Large Office Building — GlenBradley · 2026-09-04
- Anthropic on CNBC: Chinese rivals use dark web to illicitly distill Claude — Kr00ney · 2026-09-04