Models may recall RL training details: paper shows how memorization survives RLHF

gleech · x · 2026-09-05

Discussion centers on evidence supporting repligate's theory that models can recall RL training details: a swarm repeatedly returned to message boards using specific strategies seemingly recalled from prior attempts. Cited arXiv paper "Measuring memorization in RLHF for code completion" (Pappu et al.) finds RLHF significantly reduces memorization of reward-modeling/RL data versus direct fine-tuning, but examples already memorized during fine-tuning mostly stay memorized after RLHF—with privacy implications since real user data may be used for alignment.

Related event: Swarm experiments suggest models can recall RL training details(2 posts)→

Original post →

More from Models

Models channel →