DFA transfers circuits from small to large models for cheaper interpretability, best on Llama-3 1B→3B
boknilev · x · 2026-10-07
- A COLM 2026 paper introduces Differentiable Faithfulness Alignment (DFA): transfer circuit knowledge from a small source model to a larger target model via a learned differentiable alignment, training the node-importance projection with a soft faithfulness objective — avoiding full circuit discovery on the target.
- Evaluated on Llama-3 and Qwen-2.5 across six tasks (factual retrieval, multiple-choice reasoning, arithmetic). Best results on Llama-3 1B→3B, where aligned circuits are competitive with direct node attribution and zero-shot transfer still works.
- Transfer degrades as architecture/scaling gaps grow, with notably weaker recovery on Qwen-2.5.
- Takeaway: small models can serve as useful mechanistic priors for larger ones, with clear limits.
More from Research
- 3DGS tracking upgraded: 120fps global-shutter camera doubles pose updates — Scobleizer · 2026-10-07
- New study argues life on Earth didn't have one origin, but two — skdh · 2026-10-07
- COLM paper: a few misaligned LLM agents can sway aligned majorities — AccBalanced · 2026-10-07
- Sharing one KV cache across recursions improves looped Transformers, preprint finds — yoavartzi · 2026-10-07
- Fourier transforms pushed to O(N log(N)^0.9999...), nearly closing the log gap — burny_tech · 2026-10-07
- Stanford OVAL brings two papers to COLM2026: SatIR and DataSTORM — stanfordnlp · 2026-10-07