Milestone: First Double-Blind Evaluation of Proprietary Model in Secure Enclave

HaydnBelfield · x · 2026-08-27

AVERI, in collaboration with Google DeepMind, OpenMined, and MLCommons, announced a historic milestone: the first double-blind evaluation of a proprietary language model. Gemini 2.5 Flash-Lite was tested using MLCommons' AILuminate safety benchmark within a secure enclave, a hardware isolation technique protecting sensitive computations. This collaboration addresses the structural challenge in high-stakes evaluation where developers protect model weights while evaluators protect test prompts.

Related event: DeepMind and partners complete first double-blind evaluation of a proprietary frontier model(10 posts)→

Original post →

More from Safety

Safety channel →