DARLING: RL Method Balancing Quality and Diversity in LLM Responses Accepted by NeurIPS

Jason Weston's team and JHU researchers introduced DARLING, a method that jointly optimizes response quality and diversity during online RL to combat repetitive outputs from post-training. The paper has been accepted by NeurIPS.

2026-09-25 ~ 2026-09-25 · 2 related posts