TREK: Distillation-Driven Exploration Boosts Math & Agent Tasks

Yuanda Xu · hf · 2026-07-08

TREK (Distill to Explore, Reinforce to Refine) is a novel policy optimization framework. Its core innovation lies in applying knowledge distillation to the exploration phase rather than traditional imitation, breaking through the exploration bottlenecks of existing methods by expanding the policy search space, before leveraging reinforcement learning to boost final performance during the refinement phase. TREK achieves significant performance gains on challenging mathematical reasoning and agentic tasks, providing new insights into the exploration-exploitation trade-off in LLM training.

Related event: TREK Enhances GRPO via Distillation-Driven Exploration(2 posts)→

Original post →

More from Research

Research channel →