Embedded evaluator commitments by OpenAI and Anthropic are welcome but insufficient, argues Daniel Tan

MariusHobbhahn · x · 2026-09-20

Daniel Tan writes on LessWrong about the "embedded evaluators" that OpenAI and Anthropic have voluntarily committed to under their "pace the frontier" pledges. He calls it a great step but insufficient: misaligned transcripts are good evidence of misaligned models, yet the converse fails — aligned transcripts alone don't prove alignment, since aligned behavior could just mean post-training hill-climbed on your specific evaluation.

He cites Ryan Greenblatt's remark that misaligned behaviors dropping from high rates to zero looks like whack-a-mole rather than solving underlying misaligned drives. Tan argues third-party training run assessments are still needed, and that evaluation data is strongest only when confidently out-of-distribution from training data.

Original post →

More from AGI Musings

AGI Musings channel →