Psychometric methods reveal major weaknesses in AI safety benchmarks
The Decoder · rss · 2026-08-22
Researchers at the UK AI Security Institute used psychometric methods to show that popular safety benchmarks for LLMs do not measure a consistent trait. The study indicates that blanket blocking of requests can artificially inflate safety scores even as model utility decreases. It also proposes a method to catch models that act more cautious during tests than in normal use.
More from Safety
- Sam Altman on the AI dilemma: trade-offs between loss of control and power centralization — r0ck3t23 · 2026-08-24
- Debate erupts over lethal military robots vs. failing civilian units — teortaxesTex · 2026-08-24
- Only 1 of 20 Potential Presidential Candidates Answered AI Pause Query — DavidSKrueger · 2026-08-24
- Chinese Transforming Robot Dog Sparks US Trade Policy Criticism — TinfoilTricorn · 2026-08-24
- Turkey blocks at least 12 Grok posts on national security grounds — Unusual_Variation293 · 2026-08-24
- Nature Comment: Provenance, not interpretability, grounds trust in autonomous science — gabepgomes · 2026-08-24