Stella Biderman: Embedded evaluators can't fix companies that choose to be bad

BlancheMinerva · x · 2026-10-04

Stella Biderman pushes back on Dario Amodei's proposal to embed third-party evaluators in frontier AI labs. Her argument: embedded evaluators are like bank auditors or IAEA inspectors — they observe and report but need real government power to matter, which she sees as implausible. In a thought experiment, an evaluator with full access at OpenAI would log unaddressed critical risks (no air-gapped training for cyberwarfare-capable models, no chain-of-thought monitoring in pre-deployment testing). Her conclusion: the problem is companies choosing to be bad, and evaluators can't fix deliberate organizational decisions.

Original post →

More from Companies & People

Companies & People channel →