Report details first double-blind eval of LLM using secure enclaves

Miles_Brundage · x · 2026-08-27

AVERI released a pilot report detailing the world's first double-blind evaluation of a proprietary language model, Gemini 2.5 Flash-Lite, in collaboration with Google DeepMind, OpenMined, and MLCommons. The evaluation used a secure enclave and Trusted Execution Environment (TEE) to ensure model weights and test prompts remained hidden from each other, addressing benchmark contamination and structural issues in independent auditing.

Related event: DeepMind and partners complete first double-blind evaluation of a proprietary frontier model(10 posts)→

Original post →

More from Safety

Safety channel →