Neel Nanda: embedded evaluators are a great first step, but we need enforceable agreements
NeelNanda5 · x · 2026-09-13
- DeepMind researcher Neel Nanda responds to the labs' embedded evaluator commitments, calling them a fantastic first step.
- But he cautions that without an agreement to enforce them, embedded evaluators risk being empty words — "we need to see more here."
Related event: DeepMind's Nanda: Unpaced AI Progress Is Unsafe(4 posts)→
More from Safety
- AVERI Paper on Frontier AI Auditing Backs Dario's Embedded Evaluator Commitment — Miles_Brundage · 2026-09-13
- Sam Altman Echoes Dario's Pacing Call: OpenAI Will Also Grant Independent Evaluators Access — AndrewCritchPhD · 2026-09-13
- Zombie AIs: security talk on compromised AI agents that keep executing malicious instructions — wunderwuzzi23 · 2026-09-13
- Dario Amodei Calls to Pace the Frontier; Anthropic Grants Third-Party Evaluators Employee-Level Access — peterwildeford · 2026-09-13
- Andreessen: rogue AI cyber attacks smell like false flags, a convenient excuse for regulatory capture — beffjezos · 2026-09-13
- Sam Altman: stopping AI would mean kids dying of curable diseases — firstadopter · 2026-09-13