AI Security Institute tests lie detectors across 31 open-weight models
geoffreyirving · x · 2026-07-22
The AI Security Institute paper “Did you lie?” studies how to evaluate lie detectors across model scale and belief-verified model organisms.
- The authors argue that robust lie detection needs testbeds where a model can be shown to believe the opposite of what it says.
- They build on 13 reasoning-model organisms plus prompted-lying testbeds.
- Four detectors are evaluated: a chain-of-thought judge, a logprob classifier, and two activation probes, including Did-You-Lie (DYL).
- On prompted lying across 31 open-weight models, all four detectors show positive scaling with model capability.
- But every activation- and logprob-based detector degrades sharply on trained model organisms; only the chain-of-thought judge remains strong, partly because the verification process favors CoT-readable beliefs.
- The authors conclude that current lie detectors cannot support high-confidence claims about model beliefs and release the datasets, model organisms, and trained detectors.
More from Safety
- ExpSec Releases Automated Multi-Lingual Red-Teaming System for LLMs — maksym_andr · 2026-07-23
- AI safety debate: why voluntary incident disclosures still deserve praise — RyanGreenblatt · 2026-07-22
- Maintainer says an AI tool filed a high-severity report for a basic memory bug — jedisct1 · 2026-07-22
- New Legal Test Case Emerges Over GenAI Medical Advice — EricTopol · 2026-07-22
- OpenAI sued over claims ChatGPT gave dangerous medical advice in Florida case — Polymarket · 2026-07-22
- Reddit thread says big labs should be forced to re-benchmark shipped AI models — Solid-Wonder-1619 · 2026-07-22