One Attention Head Matters Across Five ICL Task Families: Mech Interp Paper Accepted at COLM 2026
JacobSteinhardt · x · 2026-09-23
- Vicky Hu's team announced their paper "Understanding In-context Learning of Addition via Activation Subspaces" was accepted to COLM 2026; Jacob Steinhardt reshared it.
- The work builds new tools identifying "extractor" and "aggregator" subspaces for ICL, starting with addition tasks and now generalizing to five ICL task families spanning both arithmetic and semantic tasks.
- Key finding: a single attention head can be important across all five task families, suggesting the mechanism transmitting ICL information may be more universal than task-specific pathways.
More from Research
- Yarin Gal: LLM-written experiments that fail to replicate are no different from any others — yaringal · 2026-09-23
- Yarin Gal: LLM experiments that don't replicate are just failures, and an AI arXiv could help — yaringal · 2026-09-23
- Oxford's Yarin Gal Proposes arXiv Ban LLM-Written Papers to Curb AI Slop — yaringal · 2026-09-23
- New paper proves fundamental confidence-efficiency bounds for transductive conformal prediction — _onionesque · 2026-09-23
- $1B and unlimited frontier tokens: where would you spend them to fix cybersecurity? — chrisrohlf · 2026-09-23
- Grok explains why DeepSeek picked DualPipe + ZeRO-1 over ZeRO-3 on 2048 H800s — TheZachMueller · 2026-09-23