DARLING: RL Method Balancing Quality and Diversity in LLM Responses Accepted by NeurIPS
Jason Weston's team and JHU researchers introduced DARLING, a method that jointly optimizes response quality and diversity during online RL to combat repetitive outputs from post-training. The paper has been accepted by NeurIPS.
2026-09-25 ~ 2026-09-25 · 2 related posts
- DARLING: diversity-aware RL beats standard RL on both quality and diversity — DanielKhashabi · 2026-09-25
- DARLING, a paper fixing repetitive LLM outputs after post-training, accepted at NeurIPS — DanielKhashabi · 2026-09-25