Apollo Research Lays Out Four Claims Any Scheming Safety Case Must Make
MariusHobbhahn · x · 2026-10-02
Apollo Research published a new post arguing that frontier AI developers must be able to show their models are not scheming — covertly working against them toward unintended goals. The post lays out four claims any scheming safety case must make, and the resources and access embedded evaluators need to verify them; Marius Hobbhahn shares more detail on the concrete safety claims they want to evaluate.
More from Safety
- AVERI blind-benchmarks Gemini inside OpenMined secure enclave, unseen by Google — iamtrask · 2026-10-02
- No Hat 2026 keynote: When Every Attacker Can Have a Research Team — WeldPond · 2026-10-02
- WIRED: AI companies' self-regulation accord is safety theater — nordicinst · 2026-10-02
- AC2 adds automated reward-hacking monitors, AI agent investigates cheating RL runs — ypatil125 · 2026-10-02
- Wired: Asking AI Companies to Self-Regulate Is Safety Theater — Wired AI · 2026-10-02
- You Can Do AI Safety From Second Place, Say Researchers on Open Models — joshalbrecht · 2026-10-02