GoodfireAI says direct weight edits cut a broken neuron-labeling bias from 94% to 5%

tszzl · x · 2026-07-29

GoodfireAI says a hackathon experiment trained an LLM to label its own neurons, but the RL run collapsed because 94% of labels started with “texts.”

Instead of rebuilding the data pipeline or retraining, the team edited the weights directly, cutting “texts” from 94% of labels to 5% with minimal side effects. The post frames this as a quick intervention on model internals rather than a conventional training fix.

Original post →

More from Research

Research channel →