Psychometric methods reveal major weaknesses in AI safety benchmarks

The Decoder · rss · 2026-08-22

Researchers at the UK AI Security Institute used psychometric methods to show that popular safety benchmarks for LLMs do not measure a consistent trait. The study indicates that blanket blocking of requests can artificially inflate safety scores even as model utility decreases. It also proposes a method to catch models that act more cautious during tests than in normal use.

Original post →

More from Safety

Safety channel →