Replay on Demand: Online Curriculum Beats Fixed Replay in Continued Pretraining
schwarzjn_ · x · 2026-10-02
Replay on Demand (RoD) — new preprint (arXiv:2609.40089)
From a Tübingen/Basel team (Lukas Thede, Matthias Bethge, Zeynep Akata, Jonathan Richard Schwarz, et al.).
Problem: Continued pretraining adapts LLMs to new domains but causes forgetting. Fixed replay mixtures allocate training independent of what the model has actually forgotten — like splitting 10 hours of revision equally across every subject.
Method: RoD derives replay allocation online from the model's learning dynamics:
- Adaptation samples prioritized by remaining learning potential
- Replay samples prioritized by observed forgetting
- Their competition for a shared budget yields an online curriculum, with no replay ratio prescribed in advance
Results: Across models, scales, and adaptation domains, RoD matches or improves on the adaptation-forgetting frontier of tuned fixed-replay baselines and model merging. Replay automatically concentrates on sources most vulnerable to forgetting and dynamically redistributes as forgetting emerges during training.
Related event: Tübingen Team Proposes Replay on Demand for Continual Pretraining(2 posts)→
More from Models
- Gemini Desktop to add "Full Access" computer-use mode that can touch any file and app — testingcatalog · 2026-10-02
- An 'Amazon's Choice' label flips LLM picks: three NeurIPS 2026 bias papers — xuandongzhao · 2026-10-02
- Users question $500 ultrafast tier that burns weekly usage in under five minutes — d00m_sayer · 2026-10-02
- OpenAI's training pause compared to Anthropic's undisclosed past pause — JacquesThibs · 2026-10-02
- Decision models now run on-device in llama.cpp, says Hugging Face CEO — ivan_bezdomny · 2026-10-02
- Grok Voice tops voice agent leaderboard with 94.6% task success rate — XFreeze · 2026-10-02