ACL Paper: Parallel Structures in Pre-training Data Yield In-Context Learning

JoshPurtell · x · 2026-08-26

This study investigates the origin of in-context learning (ICL) in pre-trained language models. The authors find that ICL ability relies on parallel structures in the pre-training data—pairs of phrases following similar templates within the same context window. Ablation experiments show that removing these structures reduces ICL accuracy by 51% (vs. 2% from random ablation). This effect persists even when excluding common patterns like n-gram repetitions and long-range dependencies, indicating the diversity and generality of parallel structures.

Related event: Study: LLM In-Context Learning Stems from Parallel Structures in Pretraining Data(2 posts)→

Original post →

More from Research

Research channel →