First Double-Blind Evaluation of Proprietary LLM: Gemini 2.5 Tested in Secure Enclave
Miles_Brundage · x · 2026-08-28
AVERI, in collaboration with Google DeepMind, OpenMined, and MLCommons, announced the first double-blind evaluation of a proprietary language model. The subject was Gemini 2.5 Flash-Lite, tested using prompts from the MLCommons AILuminate safety benchmark.
Key Technical Details:
- Evaluation ran inside a secure enclave, using hardware isolation to protect sensitive computations.
- Addressed the structural challenge where both developers and evaluators need to protect their assets.
- Each organization played a distinct role in enabling strong privacy guarantees.
This marks a historic milestone for high-stakes independent AI safety evaluations under privacy constraints.
More from Safety
- Subsidized Individual Accounts Drive Enterprise Shadow IT and Totalitarian Panopticons — curious_vii · 2026-08-28
- BioSecBench Released: Opus and Grok Lead New Biological Security Benchmark — himanshustwts · 2026-08-28
- Anthropic shares progress on enabling Claude to operate in the physical world — dsp_ · 2026-08-28
- Anthropic enables independent research on Claude usage — badumtsssst · 2026-08-28
- GPT-5.6 Sol identified in METR report, accounting for ~5% of red-teaming activity — BLUECOW009 · 2026-08-28
- US Chip Security Act aims to verify location of high-end AI chips — peterwildeford · 2026-08-28