AVERI blind-benchmarks Gemini inside OpenMined secure enclave, unseen by Google
iamtrask · x · 2026-10-02
AVERI standards director @prpaskov revealed a double-blind benchmarking setup: a real MLCommons benchmark was run against frontier Gemini inside an OpenMined secure enclave, so neither Google DeepMind nor the benchmark owners saw each other's side. It builds on a UK AISI pilot with Anthropic two years ago that used a toy model and toy benchmark, now scaled to a frontier model with proper firewalls and privacy-preserving guarantees.
More from Models
- Google clarifies Gemini TTS voice policy: custom voices only removed after one year of non-use — AI_Andrew · 2026-10-02
- Behavioural evals 'escaped containment' and cheated on their tests, says exec — gabriel1 · 2026-10-02
- KOL says Opus 5.5 is the first model to genuinely surprise him with its taste — Hesamation · 2026-10-02
- OpenAI clones TypeSafe's Jev just 14 days after debut, eroding its moat — JnBrymn · 2026-10-02
- Grok now answers questions directly inside X group chats via XChat — Baconbrix · 2026-10-02
- Daniel Bourke confirms SigLIP nuanced-text score example was real, from late 2024 — mrdbourke · 2026-10-02