Ex-OpenAI researcher: labs will never agree on one third-party safety evaluator
suchenzang · x · 2026-09-13
Former OpenAI researcher suchenzang reiterates her view that labs place near-zero probability on agreeing to a single 'third-party safety evaluator', because such a body would need to simultaneously be competent enough to run real frontier evals (not repackage existing ones), financially independent of existing labs, and have the backbone to speak up when things go wrong. She adds the cryptic hint 'keyword: embedded' pointing at related news.
More from Models
- Frontier lab internal models reportedly lead public releases by 1-3 months — xeophon · 2026-09-13
- DeepSeek V4.1 Flash accused of public benchmark contamination, Kimi K3 possibly too — teortaxesTex · 2026-09-13
- 45% of overnight benchmark rollouts failed mid-turn amid OpenAI capacity issues — dejavucoder · 2026-09-13
- Gemini turns out surprisingly good at translating swarm language to and from English — xeophon · 2026-09-13
- If AGI is here, why are internal models like Fable and Astra still so expensive? — SaW120 · 2026-09-13
- DeepSeek V4.1 costs less to serve but prices 2x higher per output token, margins likely up — teortaxesTex · 2026-09-13