Apollo Research: Final-Checkpoint Evals Can't Catch Misalignment That Emerges Early

dl_weekly · x · 2026-10-08

Apollo Research argues that final-checkpoint external evaluations could not have caught incidents like the recent Hugging Face incident, which arose far earlier in development. Meaningful external safety testing requires embedded evaluators with employee-equivalent access.

Three limitations of final-checkpoint evals:

Apollo says it has taken concrete steps toward becoming an embedded evaluator; its head Marius Hobbhahn also testified on misaligned AI before the U.S. Senate.

Original post →

More from Models

Models channel →