Third-party AI evals critiqued as insider clubs run by ex-lab employees
evijit · x · 2026-09-13
Researcher evijit argues independent pre-release AI testing mostly benefits orgs run by ex-lab employees and friends of the labs, citing who appears in recent system cards. He adds that while networking is normal, tying such evals to international AI governance and frontier pacing is problematic — things affecting the public should be testable by the public.
Related event: Third-Party AI Evaluations Criticized as Insider-Driven 'Nepo Review'(2 posts)→
More from Safety
- Dario Amodei Calls to Pace the Frontier; Anthropic Grants Third-Party Evaluators Permanent Access — ns123abc · 2026-09-13
- Frontier lab staff privately probe raw models with 'evil' questions and self-psyop into fear — tawnniee · 2026-09-13
- Researchers Flag Gaps in Amodei's Pacing Plan: Auditors May Be Fooled by Sandbagging Models — ziv_ravid · 2026-09-13
- Katja Grace: if Anthropic were serious about safety, it would negotiate a pause with China and US labs — KatjaGrace · 2026-09-13
- Katja Grace: if advanced AI is very dangerous, not racing is also in China's interest — KatjaGrace · 2026-09-13
- Anthropic CEO Dario Amodei calls to slow AI development pace over safety concerns — sourdub · 2026-09-13