First double-blind evaluation of proprietary model in secure enclave

Miles_Brundage · x · 2026-08-27

AVERI announces a historic milestone: the first ever double-blind evaluation of a proprietary language model. This collaboration with Google DeepMind, OpenMined, and MLCommons tested Gemini 2.5 Flash-Lite using MLCommons' AILuminate benchmarks inside a secure enclave. Hardware isolation protects sensitive computations. The post addresses the structural problem in high-stakes evaluation where both developers and evaluators need to protect assets, and how secure enclaves offer a solution.

Related event: Google DeepMind and partners complete first double-blind evaluation of a proprietary frontier model(9 posts)→

Original post →

More from Safety

Safety channel →