Anthropic Study: Fine-Tuned Lie Detectors Fail to Generalize OOD
PandaAshwinee · x · 2026-08-24
Anthropic published an alignment science study investigating whether training a dedicated deception monitor offers benefits beyond simply prompting a capable model.
Key Findings:
- In-distribution: Fine-tuned lie detectors outperform prompted baselines.
- Out-of-distribution (OOD): The advantage mostly vanishes. Fine-tuned detectors barely beat prompted baselines, and larger prompted models often outperform them entirely.
- Larger models are generally better at detecting lies, though the trend is not monotonic.
Context:
- Reliable lie detection raises the cost of hiding misalignment but does not align the model itself.
- Prior work (e.g., Liars' Bench) showed detectors fail to generalize across lie types.
Methodology:
- Elicited on-policy lies from open-weight models across 12 settings.
- Fine-tuned models on a binary classification task (did it lie?) and evaluated generalization by training on half the lie types and testing on the others.
More from Safety
- "Model Organisms of Misalignment": a proposed new pillar of alignment research — CFGeek · 2026-08-24
- AI Agent Phished via Email, Highlights Need for Separate Identity — _AustinCalvert_ · 2026-08-24
- Critics argue Anthropic's doom marketing backfires; decentralization and open weights proposed as real safety — arthurcolle · 2026-08-24
- Proposal: Release Failed RL Checkpoints as Better 'Model Organisms' for Safety Research — CFGeek · 2026-08-24
- Amazon reportedly buying, scanning, and destroying books for AI training — ns123abc · 2026-08-24
- 78% of Organizations Lack AI Compliance While Deploying Sensitive-Data Agents — Many_Audience7660 · 2026-08-24