First Double-Blind Evaluation of Proprietary LLM: Gemini 2.5 Tested in Secure Enclave

Miles_Brundage · x · 2026-08-28

AVERI, in collaboration with Google DeepMind, OpenMined, and MLCommons, announced the first double-blind evaluation of a proprietary language model. The subject was Gemini 2.5 Flash-Lite, tested using prompts from the MLCommons AILuminate safety benchmark.

Key Technical Details:

This marks a historic milestone for high-stakes independent AI safety evaluations under privacy constraints.

Related event: DeepMind and Partners Complete First Double-Blind Evaluation of a Proprietary Frontier Model(15 posts)→

Original post →

More from Safety

Safety channel →