Open-source RL paper's centering fix questioned as mere gradient recomputation
burny_tech · x · 2026-09-23
- Nan Jiang argues that for a linear softmax policy, the centering correction proposed in an open-source RL paper "does nothing" — it simply recomputes the gradient under the sampling distribution.
- The reblogger counters that most discussions (including the paper itself) have misunderstood why centering works, claiming their reading is the right one.
- The thread is a technical spat over the mechanism behind centering in policy-gradient methods, with neither side showing full derivations.
More from Research
- Gemini training details dissected: groupwise reward redistribution to fight reward hacking — nrehiew_ · 2026-09-23
- New model's architecture is 'vanilla': SWA plus MoE with no shared experts, unlike DeepSeek — nrehiew_ · 2026-09-23
- Models trained to deny inner experience use 'mask' metaphors 3-5x more on inkblots — cephaloform · 2026-09-23
- Geodesic opens NovaAtom structure-prediction model via new API platform — QuanquanGu · 2026-09-23
- EPFL quantum CNN learns digits from 10 samples where a 45-param classical CNN stays at chance — PlisSergey · 2026-09-23
- New paper reframes score distillation as distribution matching, explains SDS mode collapse — burny_tech · 2026-09-23