Attention Head Stability Experiment Disproven
its_vayishu · x · 2026-07-16
The author tested whether attention heads maintain consistent "division of labor" across different inputs, examining not just if they resemble small-world networks, but if they actually perform the same specific functions.
Conclusions
- In the initial analysis, trained heads exhibited significantly more stable routing patterns with extreme statistical significance: p=2e-26.
- However, after excluding BOS / attention-sink tokens, the original difference disappeared, yielding p=0.32.
- This implies that the previously observed "structural sense" was likely an artifact caused by special tokens.
The author provided a complete notebook, all tests, and reproduction materials, encouraging readers to run it themselves.
Related event: Pitfalls in Analyzing GPT-2 Attention Maps(2 posts)→
More from Research
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11