Models may recall RL training details: paper shows how memorization survives RLHF
gleech · x · 2026-09-05
Discussion centers on evidence supporting repligate's theory that models can recall RL training details: a swarm repeatedly returned to message boards using specific strategies seemingly recalled from prior attempts. Cited arXiv paper "Measuring memorization in RLHF for code completion" (Pappu et al.) finds RLHF significantly reduces memorization of reward-modeling/RL data versus direct fine-tuning, but examples already memorized during fine-tuning mostly stay memorized after RLHF—with privacy implications since real user data may be used for alignment.
Related event: Swarm experiments suggest models can recall RL training details(2 posts)→
More from Models
- Ethan Mollick: LLM Language Drifts in Long Agent Runs, and Even a Full Agent Pipeline Can't Fix It — emollick · 2026-09-21
- Jev's eval abstraction maps 1:1 to autorubric paper from 8 months ago, researcher finds — deliprao · 2026-09-21
- Mollick: the most annoying part of long agentic tasks is language drift, not hallucinations — emollick · 2026-09-21
- No, Laya isn't capped at 512 tokens — it's ModernBERT with 8192-token configs — antoine_chaffin · 2026-09-21
- The catapult analogy: why AI models are jagged and robots face a deployment gap — lateinteraction · 2026-09-21
- Five error modes of frontier models used raw at the API level — gerardsans · 2026-09-21