Sparse Readout Prism explains Logit-Lens scores via sparse features instead of tokens

Matteo He · hf · 2026-09-04

Sparse Readout Prism decomposes language model readouts into sparse features to isolate readout structure from corpus-dependent lens artifacts, offering a cleaner way to interpret Logit-Lens scores.

Original post →

More from Research

Research channel →