OpenMined's PySyft splits AI evaluation into 3 roles to scale external audits
JMateosGarcia · x · 2026-09-22
Andrew Trask argues embedded AI evaluation has a scale problem: bringing many external evaluators into each AI company's office and infrastructure is expensive, slow, and risky.
OpenMined's PySyft proposes a fix by splitting evaluation into three roles: an embedded evaluator writes reusable jobs against real model assets, an internal reviewer approves them, and external researchers receive filtered outputs without ever touching underlying data or going through new approval cycles. The approach builds on 9 years of R&D and pilots with X, DeepMind, Anthropic, Google, Microsoft, LinkedIn, Reddit, the United Nations and others.
More from Safety
- ESET finds 3,000 malicious AI skills in 900K scanned; 40% of SMBs lack AI policy — TechNadu · 2026-09-22
- ESET flags 3,000 malicious AI skills out of 900,000 scanned — TechNadu · 2026-09-22
- AI detectors now work: French model output is obvious to the eye — Afinetheorem · 2026-09-22
- Who's liable when AI breaks the law? Layered accountability, not lab dodging — gerardsans · 2026-09-22
- China reportedly probing DeepSeek and Moonshot after Anthropic alleged Claude request forwarding — Hesamation · 2026-09-22
- Dario Amodei and Sam Altman to address UN Security Council at special AI meeting — Polymarket · 2026-09-22