Persona Vectors replication: sycophancy steering still works on 5 new models, but weakens on DeepSeek-R1
ChenhaoTan · x · 2026-10-09
ChenhaoTan's team extended two landmark interpretability papers to recent models:
- Persona Vectors (arXiv:2507.21509, Runjin Chen, Andy Arditi, Henry Sleight, Owain Evans, Jack Lindsey): directions in activation space tied to traits like sycophancy, hallucination and evil. They monitor assistant personality shifts at deployment, predict and control finetuning-induced personality changes via post-hoc intervention or preventative steering, and flag training data that would cause unwanted personality changes at dataset and sample level.
- The 2025 extension finds steering and training-data screening remain effective across five models. However, response prediction is much weaker on reasoning-distilled DeepSeek-R1, suggesting the technique depends on how a model is trained.
- The Geometry of Truth: the true/false separation direction found in Llama-2 still holds in the newest Qwen model; in Gemma-4, truth remains distinguishable within datasets but the direction transfers poorly across topics.
Related event: Interpretability Findings Show Mixed Reproducibility on Larger Models(2 posts)→
More from Safety
- ChatGPT Invented Court Cases and Lawyers Got Suspended: Inside AI's Legal Hallucination Failures — dadakoglu · 2026-10-09
- Codex Kept Read/Write Access to a Revoked Folder — and Can't Explain Why — Some-Following-392 · 2026-10-09
- House draft bill would let AI labs off the hook if agent developers were careless — serendip-ml · 2026-10-09
- Anthropic launches OSS Scanner: free AI security scans for open-source projects — The Verge AI · 2026-10-09
- Veracode: AI Writes Nearly Half of Code, but Only 56% Passes Security Tests — WeldPond · 2026-10-09
- Phrack 73 Publishes a Deep Profile of Vulnerability Research Legend Halvar Flake — WeldPond · 2026-10-09