GoodfireAI says direct weight edits cut a broken neuron-labeling bias from 94% to 5%
tszzl · x · 2026-07-29
GoodfireAI says a hackathon experiment trained an LLM to label its own neurons, but the RL run collapsed because 94% of labels started with “texts.”
Instead of rebuilding the data pipeline or retraining, the team edited the weights directly, cutting “texts” from 94% of labels to 5% with minimal side effects. The post frames this as a quick intervention on model internals rather than a conventional training fix.
More from Research
- A proposal to make KV cache portable across machines, data centers, and the WAN — knowrohit07 · 2026-07-29
- TransluceAI proposes oversight foundation models to catch reward hacking at scale — JacobSteinhardt · 2026-07-29
- Why tokenizers still resist end-to-end optimization despite years of pretraining — seanmcdonaldxyz · 2026-07-29
- Student builds a local coding agent on molab and beats two 7B coder baselines — S_Conradi · 2026-07-29
- Why tokenizer optimization is hard: expensive pretraining, slow feedback, and non-differentiable design — paul_cal · 2026-07-29
- TILT improves compositional text-to-image generation with a model-intrinsic reward — Debottam Dutta · 2026-07-29