Stella Biderman: Embedded evaluators can't fix companies that choose to be bad
BlancheMinerva · x · 2026-10-04
Stella Biderman pushes back on Dario Amodei's proposal to embed third-party evaluators in frontier AI labs. Her argument: embedded evaluators are like bank auditors or IAEA inspectors — they observe and report but need real government power to matter, which she sees as implausible. In a thought experiment, an evaluator with full access at OpenAI would log unaddressed critical risks (no air-gapped training for cyberwarfare-capable models, no chain-of-thought monitoring in pre-deployment testing). Her conclusion: the problem is companies choosing to be bad, and evaluators can't fix deliberate organizational decisions.
More from Companies & People
- FAI launches two-week SF headhunting tour to recruit AI policy talent from frontier labs — KeeganMcB · 2026-10-04
- Musk: Robotaxi autonomous safety gets same rigor as Tesla's safest human-driven cars — elonmusk · 2026-10-04
- Hyperscaler veterans are jumping to Neo clouds to bet on the AI buildout — abhiadesai · 2026-10-04
- Speech AI pioneer Shinji Watanabe's Google Scholar h-index hits 100 — shinjiw_at_cmu · 2026-10-04
- Researcher gets 10 ICRA/RAL review requests in a week as robotics papers flood in — ChongZzZhang · 2026-10-04
- LessWrong guide shares unconventional advice for applying to AI safety fellowships — austinc3301 · 2026-10-04