Replay on Demand paper allocates replay data online from model forgetting dynamics
schwarzjn_ · x · 2026-10-02
A new arXiv paper from Jonathan Richard Schwarz's team introduces Replay on Demand (RoD), tackling the adaptation-vs-forgetting trade-off in continued pretraining.
- Core idea: Fixed replay mixtures ignore what the model actually needs to retain. RoD instead prioritizes adaptation samples by remaining learning potential and replay samples by observed forgetting, letting them compete for a shared training budget to yield an online curriculum.
- Results: Across models, scales, and adaptation domains, RoD matches or beats tuned fixed-replay baselines and model merging on the adaptation-forgetting frontier, without prescribing replay allocations in advance.
- Findings: Replay automatically concentrates on sources vulnerable to forgetting and dynamically grows and redistributes as forgetting emerges during training.
- The authors describe it as a simple, intuitive extension of RHO selection; learned curricula transfer across scales, amortizing cost.
Related event: Tübingen Team Proposes Replay on Demand for Continual Pretraining(2 posts)→
More from Research
- Arena's reward recipe lifts post-trained FLUX.2-dev by 69 Elo on live T2I leaderboard — arena · 2026-10-02
- AxiomicLabs' Tiny Theory of Mind benchmark hits Hugging Face's front page — Megneous · 2026-10-02
- ScholarCatalyst: New Benchmark Shows Agentic Search Loses to Plain Embedding Retrieval — lateinteraction · 2026-10-02
- Deep Learning Weekly #475: GPT-6.1 Sol, LLM-as-a-Judge, JIT memory for agents — skdh · 2026-10-02
- VASC: training-free sparse attention speeds up 3D reconstruction inference by up to 2.29x — zhenjun_zhao · 2026-10-02
- ARROW: arbitrary reconstruction and tracking of 4D observations in the wild — zhenjun_zhao · 2026-10-02