Neel Nanda: embedded evaluators are a first step, enforceable agreements needed
NeelNanda5 · x · 2026-09-13
DeepMind alignment researcher Neel Nanda argues that empty words or embedded evaluators without an enforceable agreement aren't enough: third-party evaluators are a fantastic first step, but more concrete commitments are needed.
Related event: DeepMind's Nanda: Unpaced AI Progress Is Unsafe(4 posts)→
More from Safety
- Thom Wolf agrees with 75% of Amodei's letter, but calls its lead-widening framing counterproductive — beffjezos · 2026-09-13
- e/acc founder: AI capability centralization is the opposite of safety — beffjezos · 2026-09-13
- AVERI Paper on Frontier AI Auditing Backs Dario's Embedded Evaluator Commitment — Miles_Brundage · 2026-09-13
- Zombie AIs: security talk on compromised AI agents that keep executing malicious instructions — wunderwuzzi23 · 2026-09-13
- Andreessen: rogue AI cyber attacks smell like false flags, a convenient excuse for regulatory capture — beffjezos · 2026-09-13
- Sam Altman: stopping AI would mean kids dying of curable diseases — firstadopter · 2026-09-13