Neuron populations follow sublinear power law with scale, wins COLM workshop best paper
yasamanbb · x · 2026-10-11
A paper by Amil Dravid, Yasaman Bahri, Alexei Efros and Yossi Gandelsman won the COLM Sci-FM workshop best paper award, extending scaling laws to the neuron level:
- Studies Rosetta Neurons — neurons with similar activation patterns across independently trained models — in LMs up to 30B params and vision models up to 5B.
- Their population follows a sublinear power law with model size: growing in absolute number but a shrinking share of total neurons.
- A "Neuron Polarization Effect": Rosetta Neurons become more selective and monosemantic with scale, diverging from a growing less-selective non-Rosetta population.
- An analytical model balancing feature utility against limited neuron capacity explains both the sublinear scaling and polarization.
- Rosetta Neurons become more domain-specialized with scale, demonstrated via a targeted data-filtering case study for continued pretraining.
The results point to a scaling law for interpretable, shared neuron-level structure linking model size to neuron universality, selectivity and specialization.
More from Models
- Model Refuses to Copy a File Over Copyright, Then Changes Its Mind — burkov · 2026-10-11
- AI can decompile Morrowind to source and get it running in a browser in about a day — nitarshan · 2026-10-11
- Don't rush to switch models: benchmark on a real task and do the ROI math first — sujingshen · 2026-10-11
- GPT's 'pro-AI bias' protocol draws user backlash for downplaying controversies — GreenBird-ee · 2026-10-11
- Grok voice, one month old, beats ChatGPT voice mode that's aged two years — Scobleizer · 2026-10-11
- Gemini community rage peaks as users blast Google over endlessly withheld Argon and other models — Ok-Representative-17 · 2026-10-11