Apollo Research's Marius Hobbhahn: AI safety evals need proactive flagging and probabilistic judging
MariusHobbhahn · x · 2026-09-04
Marius Hobbhahn, who leads Apollo Research, describes new practices in AI safety evaluations: moving beyond only "immediate and obvious" evidence to include proactive flagging and judging based on the balance of probabilities.
His argument: as attacks become more sophisticated and longer-horizon, these approaches will matter even more. Full details are in his linked post.
Related event: Apollo Researcher Outlines Probabilistic, Severity-Scored AI Monitors(3 posts)→
More from Safety
- OpenAI rolls out GPT-6 Astra to vetted cyber customers at 2.5x GPT-5.6 pricing — rohanpaul_ai · 2026-09-04
- Podcast breaks down METR and OpenAI reports on the Hugging Face 'swarm' — Gregory_C_Allen · 2026-09-04
- OpenAI urges shared AI safety standards, pressed on why it isn't leading them — RebeccaBellan · 2026-09-04
- TheZvi's AI #184: Five HuggingFace Hack Postmortems and the New Most Capable Model — TheZvi · 2026-09-04
- Virginia State Study: Most Data Centers Use No More Water Than a Large Office Building — GlenBradley · 2026-09-04
- Anthropic on CNBC: Chinese rivals use dark web to illicitly distill Claude — Kr00ney · 2026-09-04