Sparse Reward Subsystem in Large Language Models
Guowei Xu, Mert Yuksekgonul, James Zou
cs.CL
2026-02-01
Tsinghua and Stanford find that under 1% of neurons encode state value and step-level TD error, and zeroing them drops Qwen-2.5-7B from 75.2% to 20.3% on MATH500.
Prior probes already read whether an answer will be right from an LLM's full hidden state. The signal is decodable, and which coordinates hold it is not. If reward lives in a few units, confidence scores and process rewards can ignore most of the vector.
An autoregressive LLM is treated as a policy over an unfinished reasoning trace. The value of a prefix, V(st), is the probability that sampling onward ends in a correct answer. In maximum-entropy RL, the log probability of an optimal action scales with the advantage, so a strong policy and a value function come as a pair. A binary correct or incorrect label is one scalar for Bayes-optimal prediction. The superposition hypothesis then expects a feature used all through a trace to land on a sparse set of neurons. Both arguments are labeled as motivation.
The procedure is to train a probe, then prune it. A two-layer MLP reads the hidden state. Input units are ranked by the L1 norm of the first-layer weights, and the small weights are cut. The value probe has hidden width 1024. The dopamine probe has width 32.
Value neurons train on squared temporal-difference (TD) error. Intermediate steps target the gap between neighboring value predictions, and the last step targets the true correct or incorrect label. The discount is 1-1e-5, so discounting barely happens. Evaluation uses s0, the state after the question and before any answer token, and scores AUC against the final reward. Units that still predict once fewer than 1% remain are the value neurons.
Dopamine neurons take their name from biological cells that encode reward prediction error. They read the step-level version of that error. The answer is split into paragraphs. Monte Carlo continuations estimate value at each boundary, and the target is the discounted jump. Training and evaluation keep only paragraphs whose absolute jump exceeds 0.3. Hidden states are standardized within the response and averaged over the paragraph. The metric is Spearman rank correlation with the Monte Carlo jump.
On Qwen-2.5-7B-SimpleRL-Zoo, the top 1% of value-neuron activations in a single layer are zeroed. Controls zero the same fraction at random, by MLP down-projection magnitude, by Wanda (a pruning score that mixes weight and activation), or by a probe trained to predict the next token rather than reward.
On layers 2 to 4 of Qwen-2.5-14B-SimpleRL-Zoo, AUC on GSM8K and MATH500 barely falls as units are pruned. The same pattern appears on MATH500, ARC, the coding set MBPP+, and the instruction set IFEval, and on Qwen3.5-0.8B, Phi-3.5-mini, and Llama-3.1-8B-Instruct. Absolute AUC at s0 is only moderate. At the last token of the finished answer, AUC on Minerva Math stays mostly above 0.8 across four models.
Accuracy on MATH500 before any zeroing is 75.2%.
| Layer | Value neurons | Random | Next-token probe | Magnitude | Wanda |
| 2 | 37.0 | 77.0 | 76.4 | 74.6 | 74.8 |
| 3 | 13.6 | 73.4 | 76.8 | 61.6 | 48.2 |
| 4 | 29.4 | 73.8 | 77.4 | 75.4 | 67.0 |
| 5 | 1.2 | 74.4 | 75.2 | 75.0 | 73.4 |
| Avg | 20.3 | 74.6 | 76.4 | 71.6 | 65.8 |
Value neurons cost 54.9 points on average. Random costs 0.6, the next-token probe gains 1.2, magnitude costs 3.6, and Wanda costs 9.4. Layer 5 falls from 75.2% to 1.2%. Overlap with next-token neurons is 0.7%, against 0.5% for a random draw. A probe trained on the final reward instead of TD error, zeroed at the same 1%, only lowers accuracy to 66.4%.
At layer 3 of the 14B model, the intersection-over-union of value neurons across GSM8K, MATH500, and ARC beats a random baseline, and the overlap rises as only the highest-weighted units remain. SimpleRL-Zoo and PPO-Zero, two verifiable-reward fine-tunes of the same base, also overlap above chance. Across model families the paper shows that the sparse set exists. It does not show that the coordinates match.
Dopamine probes on layers 20 to 22 of the 7B and layers 15 to 17 of the 14B keep a flat Spearman curve over a wide pruning range. The text gives no summary correlation. The figure axis stops at 0.4. Neuron 1517 in layer 5 spikes on a key step and dips where the trace goes wrong. Zeroing the top 20% of earlier value neurons shifts those peaks and troughs. Zeroing a random 20% leaves the curve almost unchanged.
For confidence, value neurons at the best pruning ratio, averaged over layers 2 to 4, score AUC 0.67 across four models on MATH500 and ARC. LCD, a linear probe on the full hidden state, scores 0.60. Verbalized confidence scores 0.52, next-token log probability 0.48, and question length 0.62. The average ranks first. Individual cells do not: on Qwen3.5-0.8B and MATH500, question length scores 0.77 and value neurons 0.71.
For process reward, the same 7B model runs paragraph-level search on the MATH500 validation split. It draws 4 candidates per paragraph and keeps the one the dopamine probe scores highest. Mean accuracy over three seeds is 77.8%. Greedy decoding and random choice among the four candidates both score 72.2%. An implicit process reward, log probability under the RL model minus log probability under the base model, scores 75.0%. No external trained process reward model is compared.
The reward signal sits on a small set of units that can be named, and zeroing them hurts reasoning. Early value units also sit upstream of the dopamine units. The value readout can score confidence before generation starts. The dopamine readout can score intermediate steps. With 4 candidates it beats the implicit reward by 2.8 points and greedy decoding by 5.6.
The improvement is incremental. Confidence beats question length by 0.05 AUC and the full-state linear probe by 0.07. Locating the neurons still takes traces labeled right or wrong.
The paper states two limits. The method needs a reward from the environment, and open-ended generation was not tried. Models above 32B were not checked. The largest model in the reported runs is 14B.
The names are an analogy. In biology, value neurons encode subjective value and dopamine neurons encode reward prediction error. Here the names mark hidden units whose activations line up with value or TD error. The experiments do not claim that the model feels reward.
A collapse after zeroing 1% of units has another reading. Those coordinates may be ordinary early-layer computation, with reward as a correlated side readout. The next-token probe does almost no damage, and Wanda costs 9.4 points on average, which narrows that reading. Wanda alone costs 27 points at layer 3, so structural importance and reward specificity are still tangled. Identity is defined by one probe's weights. Whether a different architecture would select other units is not reported. Overlap with the next-token set is only 0.2 percentage points above chance.
The dopamine evidence is softer. There is no summary Spearman, and the activation story is mostly neuron 1517. Steps whose absolute error is at most 0.3 are dropped before training and evaluation, so the test only covers paragraphs where value already jumped. Confidence uses the pruning ratio that scored best. Search has no external process-reward baseline. AUC before generation stays moderate.