Stella Biderman: Embedded evaluators can't fix companies that choose to be bad
BlancheMinerva · x · 2026-10-09
EleutherAI researcher Stella Biderman published a critique of Dario Amodei's viral proposal to embed third-party evaluators inside frontier AI labs, arguing it won't accomplish anything.
Key points:
- Embedded evaluators are like bank auditors or IAEA inspectors — they watch and report but don't direct; they need significant government power behind them, which is hard to imagine happening
- Thought experiment: as an embedded evaluator at OpenAI with full access since last summer, her notes would read like someone slowly descending into madness
- Listed critical risks include: OpenAI isn't training its cyberwarfare-capable models on an air-gapped intranet, and isn't monitoring chain-of-thought during pre-deployment testing
She also says her advice on taking industry AI jobs has been "I don't think that working at OpenAI is moral" — "nobody has ever gone broke betting on OpenAI to do the immoral thing."
More from Companies & People
- Microsoft 365 Family to share Copilot AI benefits across members, but OneDrive cut to shared 2TB — tomwarren · 2026-10-09
- The OpenAI Files resurface: profit cap removal, weakened nonprofit board among four concerns — AaronBergman18 · 2026-10-09
- Blogger flags Anthropic's past use of "adversarial nation" to describe China — lxfater · 2026-10-09
- Tencent AI lead Yu Yi announces departure, citing exciting new opportunity — oran_ge · 2026-10-09
- OpenAI Responds to Open Letter from Three Recently Fired Employees — 141_1337 · 2026-10-09
- Terence Tao hosts guest post on what math students should do in the LLM era — keviv9 · 2026-10-09