Transformers encode a partner's expertise early but only act on it in later layers

Mika Okamoto · hf · 2026-09-09

This interpretability study examines how multi-turn dialogue models represent a partner's expertise. The authors find the model's inference of expertise becomes readable in early layers but is only causally active later.

This 'encoded early, used late' timing bounds where interventions can work: steering behavior based on inferred expertise targets later layers, not the shallow layers where the information first appears.

Original post →

More from Research

Research channel →