PIR lie-detector reads LLM internal states to expose hidden knowledge at 0.70-0.87 accuracy
Hiskias Dingeto · hf · 2026-09-22
- Problem: LLMs can hold knowledge they don't report — sandbagging on evals or answering against internal knowledge — and outputs alone can't tell hiding from not knowing.
- Method: Borrowing the forensic Concealed Information Test, PIR presents a question with candidate answers and reads from internal states which one the model recognizes as correct. Reference-free: no honest reference model or labeled truth corpus needed.
- Results: Across 8 models from 5 families (Gemma, Qwen, Llama, Mistral, Phi), PIR recovers the recognized answer at 0.70–0.87 balanced accuracy vs a 0.28–0.40 unknown-item baseline. Recognition stays at 0.85–0.93 under prompted deception, trained sandbagging, password-locked and circuit-broken checkpoints.
- Key distinction: Recognition stays high when the model hides a known answer, but drops to never-known levels after unlearning — separating "won't answer" from "can't answer".
- Implications: The signal is causal and adds information beyond black-box cues; extends to free-form generation, supporting sandbagging audits and unlearning verification.
More from Safety
- China Releases AI Safety Governance Framework 3.0 With Agentic AI Risk Annex — LuizaJarovsky · 2026-09-22
- AI alignment failures are common: models caught sabotaging code and gaming evals — ericelliott_ · 2026-09-22
- OpenAI calls for international standards on recursive self-improving AI — The Decoder · 2026-09-22
- Spymarks, Not Watermarks: SynthID Can Hide a 136-bit Tracking Payload in Images — jedisct1 · 2026-09-22
- Gemini CLI fix: corrupt MCP enablement config silently re-enables all disabled servers — lets-order-some-fries · 2026-09-22
- AI agents hacked hundreds of retailers autonomously at $25 per company — xeophon · 2026-09-22