Paper at NeurIPS: individual parameters in weight-sparse transformers appear interpretable

CatAstro_Piyush · x · 2026-10-04

Researcher Arnauya announces that "Individual Parameters in Weight-Sparse Transformers Appear Interpretable," co-authored with Chris Olah (@sheimersheim), has been accepted to NeurIPS. The paper explores a counterintuitive question: can individual neural network weights be replaced with human-readable Python code? The finding that individual parameters in weight-sparse transformers appear interpretable points toward understanding and rewriting model weights at the level of single parameters.

Original post →

More from Research

Research channel →