BeFOND: encoder-free sparse coding matches Gemma Scope with 1000x fewer training tokens

gottapatchemall · x · 2026-10-08

Researchers introduce BeFOND, an encoder-free iterative sparse coding model for interpreting LLM internals. Highlights: closed-form inference and learning rules unified as natural gradient flow on free energy; at 524k width it matches Gemma Scope's single-feature probing accuracy with over 1000x fewer samples (6M vs 8B SAE-training tokens); it keeps improving with SAE width while alternatives plateau; and the authors provide theory explaining why it works. Detailed 17-tweet thread included.

Original post →

More from Research

Research channel →