Stabilizing RL with a simple second-year probability trick: the covariance identity
PMinervini · x · 2026-09-19
zainhas shared how a basic undergraduate probability identity — E[AB] = E[A]E[B] + Cov(A,B) — can stabilize RL training. Reposting it, hyhieu226 quipped that most tricks that actually work in AI are 'second-year probability undergrad' level. A resonant take on how foundational math remains the most practical tool in RL.
More from Research
- Nearly 20 openjev models cataloged as hobbyist preps first community leaderboard — airesearch12 · 2026-09-19
- Mnemos engine builds synaptic-like agent memory via resonant engrams — RileyRalmuto · 2026-09-19
- Shared paper claims AI models have a concept of pain and avoid harm — skolnaja · 2026-09-19
- Steering Vectors as a Softer Way to Limit Reasoning Budgets? — maddie-lovelace · 2026-09-19
- MMLU-Redux paper finds ~6.5% of MMLU questions contain labeling errors — PMinervini · 2026-09-19
- NASA and IBM open-source lunar foundation model pretrained on 2M multimodal tiles — victormustar · 2026-09-19