Universal Attention Heads Found Across LLMs: Paper Reveals Neuron Scaling Laws

CSProfKGD · x · 2026-08-05

Nikhil Prakash utilized Goodfire's Silico platform to discover universal "Rosetta heads"—attention heads with similar attention patterns and OV circuits across 14 different language models of varying sizes and families.

This connects to a recent arXiv paper by Amil Dravid et al., Neuron Populations Exhibit Divergent Selectivity with Scale, which investigates how internal network structures evolve with scale. Key findings include:

Related event: Study Reveals Universal Attention Heads Across LLMs(2 posts)→

Original post →

More from Research

Research channel →