Attention Head Stability Experiment Disproven
its_vayishu · x · 2026-07-16
The author tested whether attention heads maintain consistent "division of labor" across different inputs, examining not just if they resemble small-world networks, but if they actually perform the same specific functions.
Conclusions
- In the initial analysis, trained heads exhibited significantly more stable routing patterns with extreme statistical significance: p=2e-26.
- However, after excluding BOS / attention-sink tokens, the original difference disappeared, yielding p=0.32.
- This implies that the previously observed "structural sense" was likely an artifact caused by special tokens.
The author provided a complete notebook, all tests, and reproduction materials, encouraging readers to run it themselves.
Related event: Pitfalls in Analyzing GPT-2 Attention Maps(2 posts)→
More from Research
- Structural ensembles beat single predictions in TCR:pMHC generalization study — quaidmorris · 2026-07-22
- Structural ensembles, not single predictions, drive robust TCR:pMHC generalization — quaidmorris · 2026-07-22
- A 3D ray plot shows how hard this Jacobian counterexample is to read — moultano · 2026-07-22
- LLM leaderboards are now often measuring the harness too, Gary Marcus warns — GaryMarcus · 2026-07-22
- New paper defines self-state attacks, showing OS defenses leave four agent-memory cases indistinguishable — Justgototheeffinmoon · 2026-07-22
- Krea 2 users recommend a two-pass Clownshark sampler setup for sharper image details — listopalafoto · 2026-07-22