Transformers encode a partner's expertise early but only act on it in later layers
Mika Okamoto · hf · 2026-09-09
This interpretability study examines how multi-turn dialogue models represent a partner's expertise. The authors find the model's inference of expertise becomes readable in early layers but is only causally active later.
This 'encoded early, used late' timing bounds where interventions can work: steering behavior based on inferred expertise targets later layers, not the shallow layers where the information first appears.
More from Research
- OpenAI claims agent swarm solved the 90-year-old Navier-Stokes Millennium Prize problem — basedjensen · 2026-09-09
- Marigold V2 retools diffusion transformers for sharper monocular depth estimation — huawei-bayerlab · 2026-09-09
- kornia-rs: Rust Low-Level 3D Vision Library Outperforms OpenCV and NVIDIA VPI — edgarriba · 2026-09-09
- Researcher cites more blogposts and HF links than papers, blames academia — antoine_chaffin · 2026-09-09
- ECCV 2026 workshop paper improves FoundationPose by enforcing a coherent scene — ducha_aiki · 2026-09-09
- ECCV 2026 keynote preview: Christian Rupprecht on 'Are we learning to rediscover geometry?' — ducha_aiki · 2026-09-09