DARLING: diversity-aware RL beats standard RL on both quality and diversity
DanielKhashabi · x · 2026-09-25
- Jason Weston's team introduces DARLING (Diversity Aware RL), tackling the repetition problem amplified by post-training by jointly optimizing response diversity and quality in online RL.
- The method uses a learned partition function to balance quality and diversity.
- It outperforms standard RL on both axes (e.g. higher pass@1/pass@k) and works for both verifiable and non-verifiable tasks.
More from Models
- Reported $500 ChatGPT tier likely to sell out as power users outgrow the $200 Pro plan — tengyanAI · 2026-09-25
- Early praise: OpenAI's Astra called the most thorough code review model yet — justalexoki · 2026-09-25
- ChatGPT bug: final prompt at context limit silently overwrites previous state, corrupting handoffs — xtraa · 2026-09-25
- DevDay preview: expect tools and first hardware demo, not a big model leap — haider1 · 2026-09-25
- Irregular admits AI eval incidents were environment flaws, not rogue AI behavior — robleclerc · 2026-09-25
- Cerebras plan is $500; the $600 shown includes VAT, says leaker — testingcatalog · 2026-09-25