Stanford & Tsinghua Find Two Neuron Types Encoding Reward in LLMs
Researchers from Stanford and Tsinghua identify sparse value and dopamine neurons in LLMs that encode state value and temporal difference errors, forming a causal reward subsystem enabling confidence estimation and process reward modeling for improved reasoning.
Discovery of a Reward Subsystem in Large Language Models
A joint study by Stanford University and Tsinghua University (arXiv:2602.00986) reveals that reward-related information in LLMs is not diffusely distributed but concentrated in a small subset of neurons, termed the reward subsystem . The paper identifies two distinct neuron types: value neurons and dopamine neurons .
Value Neurons: Encoding State Value
Value neurons represent the expected value of the current reasoning state — the probability of eventually producing a correct answer. Researchers trained a value probe on hidden states, then iteratively pruned input neurons. Even when retaining less than 1% of neurons , the probe accurately predicted state value. This sparsity held across models (Qwen-2.5-14B, Qwen, Phi, Llama) and tasks (GSM8K, MATH500, ARC, MBPP+, IFEval).
Causal intervention on Qwen-2.5-7B-SimpleRL-Zoo confirmed functional importance: zeroing the top 1% value neurons in a single layer dropped MATH500 accuracy from 75.2% to 20.3% , while random zeroing of the same proportion left accuracy unchanged.
Dopamine Neurons: Encoding Temporal Difference Error
Dopamine neurons encode the step-wise temporal difference (TD) error — the change in predicted probability of final correctness after a reasoning step. Positive TD error occurs when a step increases success likelihood; negative when it decreases. These neurons show sharp peaks at key logical breakthroughs and troughs at reasoning errors.
Pruning experiments show TD error information is similarly concentrated in a sparse neuron subset.
Practical Applications
Confidence Estimation: Value neurons provide early confidence signals from the initial hidden state. On MATH500 and ARC, this method achieves average AUC 0.67 , outperforming full hidden-state probes.
Process Reward Model (PRM): Dopamine neurons score candidate reasoning steps. Generating 4 candidates per step and selecting the highest dopamine-predicted reward yields 77.8% accuracy on MATH500, beating greedy decoding (72.2%) and an implicit PRM baseline (75.0%).
Conclusion
The study demonstrates that LLMs contain a sparse, causally effective reward subsystem comprising value and dopamine neurons. These neurons can be located, intervened upon, and harnessed for confidence estimation and step-level reward modeling, offering a concrete pathway to interpret and improve LLM reasoning.
Code example
来源:ScienceAI
本文
约2000字
,建议阅读
5
分钟
斯坦福与清华大学在一项研究中出了一种更细的观察方式。Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Data Party THU
Official platform of Tsinghua Big Data Research Center, sharing the team's latest research, teaching updates, and big data news.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
