Independently trained models develop 'universal' attention heads? New research shows predictable neuron scaling

CSProfKGD · x · 2026-08-05

A tweet references @nikhil07prakash's finding that independently trained language models may develop 'universal' attention heads, inspired by @AmilDravid's discovery of 'universal neurons' (Rosetta Neurons) in independently trained neural networks.

The quoted tweet from @AmilDravid notes that scaling laws describe how loss changes with scale, but do neurons inside models change predictably too? They studied vision and language models up to 30B parameters and found systematic scaling in neuron universality, specialization, and selectivity. Paper and code are available.

Related event: Study Reveals Universal Attention Heads Across LLMs(2 posts)→

Original post →

More from Research

Research channel →