Tsinghua and Stanford locate a reward subsystem in LLMs with value and dopamine neurons
jiqizhixin · x · 2026-10-01
- A Tsinghua–Stanford study locates a "reward subsystem" inside LLMs: reward-related information is not spread across the whole hidden state but concentrated in a small subset of neurons.
- Two neuron types emerge: value neurons encode the expected value of the current state—the probability that continuing from a reasoning state yields a correct answer—while dopamine neurons encode step-by-step temporal-difference error, mirroring dopamine signals in biological reinforcement learning.
- Hidden states already carry correctness, confidence, and reward signals; this work shows how those signals are internally organized, with implications for interpretability and monitoring.
More from Research
- JHU unveils PowerSim: differentiable physics that simulates and re-renders captured 3D scenes — anand_bhattad · 2026-10-01
- New preprint traces attention heads behind LLM sycophantic agreement — xuanalogue · 2026-10-01
- Post-trained Qwen3-4B doubles stock forecast score, matches frontier LMs — MengdiWang10 · 2026-10-01
- Artificial Analysis posts full GPT-6.1 Sol evals, rolls out Intelligence Index v4.3 — ArtificialAnlys · 2026-10-01
- Edward Kmett ships Turbo Haskell: a GraalVM JIT for GHC Core that can compile GHC itself — rickasaurus · 2026-10-01
- Professor's Guide: How to Actually Understand Proofs When LLMs Do the Derivations — nanjiang_cs · 2026-10-01