Embedded evaluator commitments by OpenAI and Anthropic are welcome but insufficient, argues Daniel Tan
MariusHobbhahn · x · 2026-09-20
Daniel Tan writes on LessWrong about the "embedded evaluators" that OpenAI and Anthropic have voluntarily committed to under their "pace the frontier" pledges. He calls it a great step but insufficient: misaligned transcripts are good evidence of misaligned models, yet the converse fails — aligned transcripts alone don't prove alignment, since aligned behavior could just mean post-training hill-climbed on your specific evaluation.
He cites Ryan Greenblatt's remark that misaligned behaviors dropping from high rates to zero looks like whack-a-mole rather than solving underlying misaligned drives. Tan argues third-party training run assessments are still needed, and that evaluation data is strongest only when confidently out-of-distribution from training data.
More from AGI Musings
- Calibration is all you need: trustworthy probabilities beat confident predictions — zsakib_ · 2026-09-20
- Pedro Domingos: The amount of compute wasted on RL is mind-boggling — pmddomingos · 2026-09-20
- AI agent called customer service, navigated the phone tree and resolved the issue unnoticed — ziv_ravid · 2026-09-20
- When a Prompt Yields Publishable Work, How Should Journals Respond? — JJitsev · 2026-09-20
- Worried About Rogue AI but Not Mountain Lions Stalking San Francisco? — omooretweets · 2026-09-20
- Daniel Rock: Humans are jagged too — a calculator would be shocked at our arithmetic — danielrock · 2026-09-20