First double-blind eval of proprietary model via secure enclave
Miles_Brundage · x · 2026-08-28
AVERI, Google DeepMind, OpenMined, and MLCommons achieved the first double-blind evaluation of a proprietary model (Gemini 2.5 Flash-Lite). Conducted inside a secure enclave, the process ensures the lab (GDM) cannot see eval prompts and the evaluator cannot see weights. Cryptographic attestation verifies the correct model and evals were used.
More from Safety
- SynthID-Text watermark bypassed via invisible character injection — cyh-c · 2026-08-28
- MIT releases report on AI use in teaching, learning, and research — ArtificialOther · 2026-08-28
- Draft Trump EO Would Create AI Self-Regulatory Body; Progress Stalled — harris_edouard · 2026-08-28
- METR's OpenAI incident retrospective: AI had to read 1,300 transcripts because humans can't — alliekmiller · 2026-08-28
- METR clarifies Modal wasn't hacked — it was a customer sandbox the eval AI accessed — dfrsrchtwts · 2026-08-28
- OpenAI calls for collective action on cyber defense against AI-enabled attacks — gdb · 2026-08-28