Kuaishou's DARA Cuts Multi-Reward RL Training Steps by Up to 65%
kuaishou · hf · 2026-10-02
- Analysis: Via advantage energy (sum of squared advantages per reward), Kuaishou researchers show that under idealized GDPO normalization this energy is proportional to active-group density, exposing a residual batch-level signal imbalance across rewards.
- Method: DARA derives an inverse-square-root density correction that upweights less frequently active rewards, computing weights per rollout batch without changing the underlying policy optimization objective.
- Results: On tool calling and mathematical reasoning, DARA learns targeted behaviors faster than GDPO — up to 26% fewer steps to high format compliance and up to 65% fewer steps to near-saturated length compliance, with competitive final performance. Code released on GitHub.
More from Research
- HIDE Benchmark Exposes Memory Gaps in Robotic Manipulation Under Partial Observability — Yansong Shi · 2026-10-02
- InterEvolve Evolves Reward Programs at Test Time to Teach Humanoid Robots New Skills — UIUC-CS · 2026-10-02
- Smaller Models Make Better Rejects: Study Rethinks Preference Distillation from 7B to 72B — LinkedIn · 2026-10-02
- HeteroFold Enables Prefill-Free Cross-Family KV Cache Transfer, 10.7x Faster at 32K Context — UniversityofSouthernCalifornia · 2026-10-02
- Nature cover: 6.3M nuclei sequenced to build largest human prefrontal cortex cell atlas — jiqizhixin · 2026-10-02
- Detect LLM hallucinations in 1.3µs on CPU — but 120B models hallucinate with unanimous false certainty — More_Slide5739 · 2026-10-02