AI Evaluator Forum brings together Transluce, METR, RAND for independent AI evaluations

typewriters · x · 2026-09-11

The AI Evaluator Forum unites independent research organizations focused on rigorous technical evaluations of general-purpose AI systems in the public interest. Members include Transluce (open, scalable tech for understanding AI behaviors), METR (evaluations of frontier AI's ability to complete complex tasks without human input), RAND (dangerous-capability testing for evidence-based policy), SecureBio (biosecurity misuse risk evaluation), and the Princeton Holistic Agent Leaderboard (standardized, cost-aware, third-party agent leaderboard). The poster welcomes these members joining to extend the Forum's evaluation efforts. Membership is limited to organizations that publish rigorous independent evaluations.

Original post →

More from Safety

Safety channel →