AI Evaluator Forum brings together Transluce, METR, RAND for independent AI evaluations
typewriters · x · 2026-09-11
The AI Evaluator Forum unites independent research organizations focused on rigorous technical evaluations of general-purpose AI systems in the public interest. Members include Transluce (open, scalable tech for understanding AI behaviors), METR (evaluations of frontier AI's ability to complete complex tasks without human input), RAND (dangerous-capability testing for evidence-based policy), SecureBio (biosecurity misuse risk evaluation), and the Princeton Holistic Agent Leaderboard (standardized, cost-aware, third-party agent leaderboard). The poster welcomes these members joining to extend the Forum's evaluation efforts. Membership is limited to organizations that publish rigorous independent evaluations.
More from Safety
- New essay argues the 'rogue agent' panic conflates at least five distinct incidents — mimi10v3 · 2026-09-11
- tszzl: AI extinction risk is tiny but orders of magnitude above anything else — and better models will help alignment — deanwball · 2026-09-11
- Nearly 10% of exposed LiteLLM gateways accept default admin key 'sk-1234' — Thionne_WTZ · 2026-09-11
- Alignment's intensional definition problem: you must define 'agent', 'goals' first — zetalyrae · 2026-09-11
- Kelsey Piper: Labs plan to automate AI R&D with AI, shrinking human oversight within two years — round · 2026-09-11
- Anthropic says it stopped attempts to use models for potential biological weapons — connoraxiotes · 2026-09-11