MLP paper shows neurons turn monosemantic in clustered regression, challenging global subspace view

burkov · x · 2026-09-07

This paper challenges the common view that neural networks learn a single global low-dimensional subspace of predictive directions. For clustered data — where each cluster has its own predictive direction and response function — local prediction can be simple while the directions collectively span the full ambient dimension, leaving no useful global low-dimensional structure.

Studying standard MLPs on this clustered regression setting, the authors show the networks develop a different form of feature learning: individual neurons become monosemantic, each strongly aligning with the predictive direction of one specific cluster. The network thereby discovers both the latent cluster structure and the per-cluster prediction rules.

Original post →

More from Models

Models channel →