AI evaluator ecosystem runs wider than METR — NDA-bound orgs need regulation, not pledges
evijit · x · 2026-09-13
Responding to the METR discussion, prpaskov notes the evaluator ecosystem already runs wider than its current focus suggests: many strong evaluation orgs operate under tight NDAs and can't speak publicly despite core contributions to model cards. Voluntary commitments aren't enough, the author argues — standards and regulation are needed to bring talent into existing evaluators and seed new orgs across domains.
More from Companies & People
- Sam Altman says OpenAI is delaying its IPO, calls going public this year an "ill-advised moment" — thesaraharminta · 2026-09-13
- An open letter to Dario Amodei: if you mean it, open the weights — routelastresort · 2026-09-13
- Josh Purtell accuses frontier labs of seeking power, not safety, behind 'pace the frontier' regulation — JoshPurtell · 2026-09-13
- Grok Bot's product lead tells all: tiny team hit millions of users in seven weeks — lennysan · 2026-09-13
- 124x surge in token consumption by OpenAI researchers, flagged by Erik Brynjolfsson — erikbryn · 2026-09-13
- Reddit Skeptic: Anthropic's 'Pacing' Pledge Will Be a Few Weeks of Safety Review, Then Business as Usual — tunicamycinA · 2026-09-13