AVERI completes first double-blind evaluation of proprietary LLM

Miles_Brundage · x · 2026-08-31

AVERI, in collaboration with Google DeepMind, OpenMined, and MLCommons, announced the first double-blind evaluation of a proprietary language model (Gemini 2.5 Flash-Lite). Conducted in a secure enclave using MLCommons' AILuminate benchmarks, the initiative addresses structural privacy issues in high-stakes independent AI evaluation.

Original post →

More from Safety

Safety channel →