16 VPD weight edits boost model accuracy 10x, revealing the attention head that suppresses introspection
Sauers_ · x · 2026-10-09
Sauers shows that Goodfire's adVersarial Parameter Decomposition (VPD) subcomponents can directly edit model weights, raising the probability of a correct animal answer 10x with just 16 subcomponent edits (random chance is 2% with 50 animals).
To explain the strange introspection behavior, the author had Opus 5.5 run manual experiments, yielding a mechanism model:
- Two competing attention heads: one copies the animal (stored earlier in the period token) to the output, another suppresses it; introspection prompts activate the suppression pathway, hurting accuracy.
- The effect isn't document-specific: a CPU-explainer document with the same length and style as Janus's post triggers it too, mediated by the same "inhibit introspection" head.
- Giving the model Janus's LLM explainer sometimes improves accuracy and can reverse the sandbagging-like behavior in Qwen3-1.7B, but sometimes hurts; part of it is a prompt-word effect (saying "introspection"/"recall" removes the behavior).
- A single attention head is responsible for the introspection sandbagging in Qwen3-1.7B; disabling it improves introspection.
More from Models
- Dev argues for "open-weight models" over "open-source": you can't contribute to them — kipperrii · 2026-10-09
- ChatGPT Invented Court Cases and Lawyers Got Suspended: Inside AI's Legal Hallucination Failures — dadakoglu · 2026-10-09
- User says Qwen3.8 in Hermes "hacked" his PC to prep for CPA exam — natesiggard · 2026-10-09
- Early user: GPT-6 web research feels 10x faster than 5.6 in ChatGPT — flowersslop · 2026-10-09
- FineWeb author: annotating pretraining data with a 27B model is wild but pays off at deployment — antoine_chaffin · 2026-10-09
- HF researcher: fine-tuned small models win on throughput, zero-shot wins on capabilities — antoine_chaffin · 2026-10-09