BeFOND: encoder-free sparse coding matches Gemma Scope with 1000x fewer training tokens
gottapatchemall · x · 2026-10-08
Researchers introduce BeFOND, an encoder-free iterative sparse coding model for interpreting LLM internals. Highlights: closed-form inference and learning rules unified as natural gradient flow on free energy; at 524k width it matches Gemma Scope's single-feature probing accuracy with over 1000x fewer samples (6M vs 8B SAE-training tokens); it keeps improving with SAE width while alternatives plateau; and the authors provide theory explaining why it works. Detailed 17-tweet thread included.
More from Research
- New HSS Benchmark Shows Top AI Models Fail Basic Intuitive Visual Reasoning Humans Find Easy — dustinvtran · 2026-10-08
- Rowan Sci finds frontier LLMs surprisingly good at inverse folding, with some help — AllThingsApx · 2026-10-08
- Researchers from frontier labs debate continual learning for 3 hours — ysu_nlp · 2026-10-08
- Cambridge lab uses Google's Co-Scientist to compress 2-3 years of research into 6 months — vivnat · 2026-10-08
- CVPR 2026 ReLearn workshop: Efros argues AI may already learn too much from humans — TheZachMueller · 2026-10-08
- Researchers compile a larval fish brain into a simulated circuit running 26x faster than spiking models — mtizard · 2026-10-08