Apollo Research publishes principles for embedded evaluations of frontier AI
MariusHobbhahn · x · 2026-10-01
Frontier AI companies have committed to giving outside evaluators employee-like access to training, evaluation, and deployment. Apollo Research argues the impact hinges on implementation and has published principles for embedded evaluations — designed to incentivize safety fairly and productively for both parties. Marius Hobbhahn shared additional details, calling it a start that will keep evolving.
More from Safety
- binarybits: We Should Forecast Emerging AI Risks, Warned of AI Hacking Tools Back in 2023 — binarybits · 2026-10-01
- Senate 'Rogue AI' Hearing to Feature METR, Apollo Research and Daniel Kokotajlo — Turn_Trout · 2026-10-01
- AI Safety Researcher Turn_Trout Submits Statement to US Senate Hearing on Rogue AI Agents — Turn_Trout · 2026-10-01
- Bank of England Governor Bailey: Regulating AI 'not the right place to start', test first — alexvoica · 2026-10-01
- Gary Marcus amplifies law professor's case that OpenAI's lawyers 'have a LOT to worry about' — GaryMarcus · 2026-10-01
- FTC opens probe into OpenAI, Anthropic and eval firm METR, may compel exec testimony — ns123abc · 2026-10-01