IOI circuit findings from GPT-2 small break down in larger models, new interpretability study shows

ChenhaoTan · x · 2026-10-09

A follow-up study revisits the classic Indirect Object Identification circuit work by Wang et al., which found that 26 attention heads in GPT-2 small track repeated names and pick the right recipient in sentences like "John gave a drink to Mary".

Key findings:

Related event: Interpretability Findings Show Mixed Reproducibility on Larger Models(2 posts)→

Original post →

More from Research

Research channel →