IOI circuit findings from GPT-2 small break down in larger models, new interpretability study shows
ChenhaoTan · x · 2026-10-09
A follow-up study revisits the classic Indirect Object Identification circuit work by Wang et al., which found that 26 attention heads in GPT-2 small track repeated names and pick the right recipient in sentences like "John gave a drink to Mary".
Key findings:
- The circuit explains less of the behaviour in larger GPT-2 models and may not appear in other model families at all.
- Heads that copy the correct name remain important, but the rest of the computation spreads across more heads.
- Interpretability conclusions drawn on small models may not generalize, so cross-model validation is essential.
Related event: Interpretability Findings Show Mixed Reproducibility on Larger Models(2 posts)→
More from Research
- China Telecom and MemTensor unveil HaluMem, first operation-level benchmark for agent memory hallucinations — jiqizhixin · 2026-10-09
- BAAI's AREX research agent checks answers requirement-by-requirement, hits 82.5% BrowseComp — DeepLearningAI · 2026-10-09
- New piece lays out how to build RL environments aimed at superintelligence — JenniferHli · 2026-10-09
- Quantum Counterfactuals: Quantum RNGs as an Exploration Source for RL — jessi_cata · 2026-10-09
- OpenAI theorem drop collides with researchers' work: stronger bounds but 'unreadable' proof — guyvdb · 2026-10-09
- AI-written science floods preprint servers; researchers propose decision language models as filter — lpachter · 2026-10-09