OpenMined proposes three-role PySyft workflow to scale independent AI evaluations

iamtrask · x · 2026-09-22

OpenMined published a 32-minute research piece on scaling external AI evaluation. While Anthropic, OpenAI, and xAI have all pledged independent evaluator access and new oversight proposals emerge, two structural problems remain: embedded evaluator seats are extremely costly (security clearance, privacy review, access management), and evaluators who expose their tests risk benchmarks being gamed or leaked.

Their answer: PySyft splits evaluation into three roles — an embedded evaluator writes a reusable job against real model assets, an internal reviewer approves it, and external researchers receive filtered outputs without ever touching underlying data or going through new approval cycles.

Original post →

More from Safety

Safety channel →