ACL Paper: Parallel Structures in Pre-training Data Yield In-Context Learning
JoshPurtell · x · 2026-08-26
This study investigates the origin of in-context learning (ICL) in pre-trained language models. The authors find that ICL ability relies on parallel structures in the pre-training data—pairs of phrases following similar templates within the same context window. Ablation experiments show that removing these structures reduces ICL accuracy by 51% (vs. 2% from random ablation). This effect persists even when excluding common patterns like n-gram repetitions and long-range dependencies, indicating the diversity and generality of parallel structures.
More from Research
- TMLR Paper: Rigorous derivation of Adjoint Matching via Stochastic Maximum Principle — rishabh16_ · 2026-08-26
- CIDER Dataset: Personalized Privacy Preference Alignment — tianshi_li · 2026-08-26
- AI Analysis of Chest CTs Links Thymus Health to Lung Cancer Survival Outcomes — EricTopol · 2026-08-26
- Palomar: A Public Archive for Machine-Checked Math in the AI Era — repligate · 2026-08-26
- Presentation on weird, creative ideas for evaluating AI benchmarks — mariofilhoml · 2026-08-26
- Rasyn Lab Releases Synthon 350M for Single-Step Retrosynthesis — ycombinator · 2026-08-26