Stabilizing RL with a simple second-year probability trick: the covariance identity

PMinervini · x · 2026-09-19

zainhas shared how a basic undergraduate probability identity — E[AB] = E[A]E[B] + Cov(A,B) — can stabilize RL training. Reposting it, hyhieu226 quipped that most tricks that actually work in AI are 'second-year probability undergrad' level. A resonant take on how foundational math remains the most practical tool in RL.

Related event: Together AI Proposes Score Centering to Stabilize Off-policy RL via a Covariance Identity(5 posts)→

Original post →

More from Research

Research channel →