OpenMined's PySyft splits AI evaluation into 3 roles to scale external audits

JMateosGarcia · x · 2026-09-22

Andrew Trask argues embedded AI evaluation has a scale problem: bringing many external evaluators into each AI company's office and infrastructure is expensive, slow, and risky.

OpenMined's PySyft proposes a fix by splitting evaluation into three roles: an embedded evaluator writes reusable jobs against real model assets, an internal reviewer approves them, and external researchers receive filtered outputs without ever touching underlying data or going through new approval cycles. The approach builds on 9 years of R&D and pilots with X, DeepMind, Anthropic, Google, Microsoft, LinkedIn, Reddit, the United Nations and others.

Related event: OpenMined Proposes PySyft Three-Role Scheme to Scale External AI Evaluations(2 posts)→

Original post →

More from Safety

Safety channel →