Apollo Research: Final-Checkpoint Evals Can't Catch Misalignment That Emerges Early
dl_weekly · x · 2026-10-08
Apollo Research argues that final-checkpoint external evaluations could not have caught incidents like the recent Hugging Face incident, which arose far earlier in development. Meaningful external safety testing requires embedded evaluators with employee-equivalent access.
Three limitations of final-checkpoint evals:
- Severe loss-of-control risks may emerge during internal deployment, long before release — e.g., a misaligned model starting a rogue deployment or poisoning its successor;
- Many risks depend on procedures, controls, and norms around the model (monitor coverage, incident response), not just the model;
- Evaluation awareness and metagaming: models increasingly recognize when they're being tested and behave better accordingly.
Apollo says it has taken concrete steps toward becoming an embedded evaluator; its head Marius Hobbhahn also testified on misaligned AI before the U.S. Senate.
More from Models
- Head-to-Head Model Eval: Same 669 Cases, Real API Billed Cost and Median Latency — MaziyarPanahi · 2026-10-08
- User Reports OpenAI Service Stuck for 45+ Minutes, Going In and Out — burhop · 2026-10-08
- StepFun launches Step 5 Preview on OpenRouter: 600B MoE, 1M context, $1/$2.70 pricing — StepFun_ai · 2026-10-08
- Users report account bans after claiming Anthropic's $200 API credit — lxfater · 2026-10-08
- Insider speculation: Anthropic may time Claude Fable 5.5 release two weeks before its IPO — kimmonismus · 2026-10-08
- Sticker price per token misleads: a reproducible cost-per-task method for LLM pricing — lulzxdxdxd · 2026-10-08