Elegant Derivation Shows Policy Gradient Score Centering Is a Straight-Through Estimator
YouJiacheng · x · 2026-09-20
YouJiacheng shares a compact derivation showing that score centering in policy gradients is equivalent to a straight-through estimator (STE). With s = ∇z log p(a) = ea − p, its expectation under q is s̄ = q − p, so the centered score ea − q equals the STE form zp + sg(zq − zp). The argument applies beyond linear policies; the interpretation was first found by @ACSync, echoing @nanjiangcs's argument.
More from Research
- Podcast: How surgical data science teaches AI to understand what happens in surgery — ddonoho · 2026-09-20
- Dan Hendrycks Proposes 'Eigenism,' an Ethics Framework Making Human Flourishing AI Self-Interest — basedjensen · 2026-09-20
- Parallel structured LLM answers never check each other: the Zhaozhou MU problem — Successful-Farm5339 · 2026-09-20
- Advanced Matrix Factorization Jungle: A Living Map of Structured Factorization Algorithms and Phase Transitions — IgorCarron · 2026-09-20
- Sentence Transformers models quietly dominate Hugging Face's most-downloaded list — tomaarsen · 2026-09-20
- "Recipe for intelligence" paper published in Neuron — summerfieldlab · 2026-09-20