Google tests double-blind AI benchmarks with cryptographic protection

The Decoder · rss · 2026-08-28

Google Deepmind is piloting a double-blind evaluation for frontier AI models. Using cryptographic protection via Confidential Space, the method prevents Google from seeing test questions and evaluators from seeing model weights. The pilot with Singapore's AI Safety Institute uses Gemini Flash Lite to set a new standard for tamper-proof benchmarks.

Original post →

More from Safety

Safety channel →