Anthropic mechinterp team traced the bump in Kaplan scaling laws to emerging induction heads
gordic_aleksa · x · 2026-09-14
A recap of a finding from Anthropic's mechinterp team: the bump in Kaplan et al.'s scaling law plots was caused by the emergence of induction heads.
- An induction head is an attention head performing "[A][B] … [A] → [B]": on seeing token A, it locates a previous occurrence of A in context and copies the token that followed it
- The more abstract A and B become, the closer this gets to in-context learning rather than lexical copying
- Key result: 1-layer transformers never form induction heads and skip this phase transition, yielding less steep loss curves; emergence requires 2+ layers
Related event: Anthropic links scaling law bump to emergent induction heads(2 posts)→
More from Research
- Recurrent looped transformer offers infinite reasoning depth, code released — 141_1337 · 2026-09-15
- Researchers clash over arXiv 'oral exams' proposal amid flood of AI-written submissions — RishiBommasani · 2026-09-15
- Jay Alammar at PyData: ten questions that explain benchmark score gaps — JayAlammar · 2026-09-15
- RLE-Bench: 48 Tasks Benchmark LLM Agents on Full Robotics Engineering, Not Just Control — daibond_alpha · 2026-09-15
- BrainGPT creator: AI discoveries will soon be as explainable as quantum mechanics to a dog — every · 2026-09-15
- Ethan Mollick tests GPT on Linear A, warns it's 'researchslop' until GPT-6 cracks it — emollick · 2026-09-15