Interpretability Findings Show Mixed Reproducibility on Larger Models
ChenhaoTan's team extended two interpretability studies to newer 2025-2026 models: the IOI circuit's 26 attention heads failed to transfer, while Persona Vectors' sycophancy and hallucination directions remained monitorable and controllable.
2026-10-09 ~ 2026-10-09 · 2 related posts
- Persona Vectors replication: sycophancy steering still works on 5 new models, but weakens on DeepSeek-R1 — ChenhaoTan · 2026-10-09
- IOI circuit findings from GPT-2 small break down in larger models, new interpretability study shows — ChenhaoTan · 2026-10-09