DRIFT: A New Framework from Tsinghua and Beike for Continuous LLM Self-Evolution
青稞AI · wechat · 2026-07-26
A joint research team from Tsinghua University, Beike, and ENS Paris-Saclay introduced DRIFT, an online self-evolution post-training framework for LLMs. It aims to solve the challenge of how models can continuously improve after RL without collapsing.
The framework features several core mechanisms:
- Dynamic Difficulty Routing & Rhythm-Gated Exploration: These help the model identify critical reasoning nodes, determining when to apply RL versus self-distillation.
- Success Buffer: Combined with curriculum learning, it ensures more stable policy optimization.
Experiments show that DRIFT enables self-correction and continuous evolution without external expert supervision, achieving new SOTA on multiple complex reasoning and Tool Use benchmarks. The authors will host an upcoming webinar to dive deep into the algorithm's design and theoretical motivations.
More from Research
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- VidMap uses RoMa coarse matching on all frames, fine-scale only for keyframes — ducha_aiki · 2026-09-11
- Bug Hunt Bench author: leaderboard noise is about 2-3 points — PawelHuryn · 2026-09-11
- PNAS paper shows a tiny billiard-ball system is a universal computer — undecidability lives in two dimensions — eigensteve · 2026-09-11
- New paper: Absolute pose estimation from affine cues and gravity direction — ducha_aiki · 2026-09-11
- LoMa Paper Ships REALLY HardPairs Dataset, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11