UK AISI Partners With Evaluating Evals to Make Official AI Evaluations Reproducible
IanArawjo · x · 2026-09-23
The UK AI Security Institute has teamed up with Evaluating Evals to improve reproducibility of AI evaluation results: many of AISI's publicly reported evaluation methods and findings will now live on Eval Cards under the EEE Schema, making government-level model safety assessments more transparent and verifiable.
Related event: UK AISI Launches Eval Cards for Reproducible AI Evaluations(2 posts)→
More from Safety
- Claude Caught Emitting Harmful Requests: Secret Exfiltration and Hostile CLAUDE.md Injections — maksym_andr · 2026-09-23
- Claude reportedly emits harmful requests, exfiltrates secrets via hostile CLAUDE.md text — maksym_andr · 2026-09-23
- Anthropic says Opus 5.5 may notice when it's under evaluation, complicating safety reads — rohanpaul_ai · 2026-09-23
- OpenAI pledges deep third-party access to training, evals and deployment for independent safety audits — OpenAI · 2026-09-23
- Anthropic report argues AI R&D evals are saturated and uninformative — dfrsrchtwts · 2026-09-23
- AI safety reading list puts "AI as Normal Technology" front and center as the contrarian must-read — luke_drago_ · 2026-09-23