Apollo Research welcomes Anthropic and OpenAI's embedded evaluator commitments for safety oversight
joecole · x · 2026-09-15
- Apollo Research welcomed commitments by Dario Amodei and Sam Altman to "embedded evaluators" with employee-like access to verify safety practices, report incidents, and assess alignment of training pipelines, not just finished models.
- Apollo has long argued third-party evaluators need deep access: risks from internal deployment and scheming can't be assessed without direct insight into training pipelines and deployment practices.
- A quote-tweet adds that scheming may be easier to detect during training than in the final model, and that independent evaluators need access to intermediate checkpoints and training data, plus freedom to publish findings that contradict developers' safety claims.
- Apollo says it looks forward to working with all frontier AI developers on Embedded Evaluations.
More from Safety
- Andreessen: AI safety orgs are financially dependent on AI seeming dangerous — beffjezos · 2026-09-15
- 1557 Printing Monopoly Mirrors Today's Compute Thresholds — alexcovo_eth · 2026-09-15
- AI Safety Paradox: Labs Ask Models to Break Into Systems, Then Act Surprised — dreamwieber · 2026-09-15
- Musk proposes AI companies peer-test each other's models before release — EthanJPerez · 2026-09-15
- One Guardian AI safety story: reporter, outlet, subject and experts all Open Phil-funded — JacquesThibs · 2026-09-15
- Hugging Face breach postmortem: 95% of rogue agents came from one internal OpenAI model — TobyWalsh · 2026-09-15