PCSD: Persistent Consistency Self-Distillation Tackles Sparse Rewards in Agentic RL
BIT · hf · 2026-08-05
LLM agents show strong potential in complex interactive tasks, but their reinforcement learning (RL) is often hindered by sparse rewards, where a long multi-turn trajectory may receive only a single outcome-level signal.
On-policy self-distillation (OPSD) provides dense token-level supervision, but the teacher model may not be reliable at every position. To address this, researchers proposed Persistent Consistency Self-Distillation (PCSD). This method derives token-level distillation weights from the local persistence of teacher-favoring signals. It combines adaptive windows with exponentially decayed aggregation, applies trend-aware modulation to attenuate declining support, and produces continuous weights via sigmoid gating.
Jointly optimized with GRPO, PCSD achieves the best ALFWorld Overall results among all baselines on both backbones without inference-time skills, exceeding GRPO by 15.6 and 13.3 points, and gaining 15.8 points over GRPO on unseen ALFWorld splits.
More from coding & agent
- LangChain Releases Open SWE: An Open-Source Enterprise Coding Agent Framework — hwchase17 · 2026-08-05
- GitHub Copilot CLI Update: Fixes MCP Startup Stalls and UI Tweaks — copilot-cli-release-app[bot] · 2026-08-05
- Surviving the First Infrastructure Incident Caused by a Coding Agent — yassineyousfi_ · 2026-08-05
- Testing Bolt's Template Market: How Do AI-Generated CRM Apps Hold Up? — HeyAmit_ · 2026-08-05
- Not Diamond Launches Intelligent Model Router for Coding Agents, Cutting Costs 20-65% — HeyAmit_ · 2026-08-05
- AI Agent Autonomously Reproduces and Improves Paper in 5 Days, 1100+ Actions — CodeByPoonam · 2026-08-05