AVERI blind-benchmarks Gemini inside OpenMined secure enclave, unseen by Google

iamtrask · x · 2026-10-02

AVERI standards director @prpaskov revealed a double-blind benchmarking setup: a real MLCommons benchmark was run against frontier Gemini inside an OpenMined secure enclave, so neither Google DeepMind nor the benchmark owners saw each other's side. It builds on a UK AISI pilot with Anthropic two years ago that used a toy model and toy benchmark, now scaled to a frontier model with proper firewalls and privacy-preserving guarantees.

Original post →

More from Models

Models channel →