DRIFT: A New Framework from Tsinghua and Beike for Continuous LLM Self-Evolution
青稞AI · wechat · 2026-07-26
A joint research team from Tsinghua University, Beike, and ENS Paris-Saclay introduced DRIFT, an online self-evolution post-training framework for LLMs. It aims to solve the challenge of how models can continuously improve after RL without collapsing.
The framework features several core mechanisms:
- Dynamic Difficulty Routing & Rhythm-Gated Exploration: These help the model identify critical reasoning nodes, determining when to apply RL versus self-distillation.
- Success Buffer: Combined with curriculum learning, it ensures more stable policy optimization.
Experiments show that DRIFT enables self-correction and continuous evolution without external expert supervision, achieving new SOTA on multiple complex reasoning and Tool Use benchmarks. The authors will host an upcoming webinar to dive deep into the algorithm's design and theoretical motivations.
More from Research
- ICML 2026 oral paper replication scores stay middling after a stricter re-scoring — profjamesevans · 2026-07-27
- Long-running agents will need immutable event logs, this thread argues — sebpaquet · 2026-07-27
- Seed IQ navigates Doom II, prompting questions about benchmarks beyond ARC-AGI — Fit_Transition8824 · 2026-07-27
- Agentic Data Science in Practice: Agents Write Code but Answer Wrong Questions — hugobowne · 2026-07-27
- A concise canon of foundational papers in ML, systems, NLP, speech, and audio — deliprao · 2026-07-27
- TechCrunch says brain-wave signals could be the next unlock for physical AI training — TechCrunch AI · 2026-07-27