MoDA: RL alignment method fights LLM mode collapse while preserving output quality
stanfordnlp · x · 2026-09-19
A Stanford-led paper (arXiv:2609.14896, authors include Yejin Choi and Natasha Jaques) introduces MoDA (Mode-conditioned Diversity Alignment), targeting the mode collapse that alignment training inflicts on LLM output diversity.
Key points:
- Inspired by coordination in multi-agent RL, MoDA runs online post-training RL on a single shared LLM policy conditioned on abstract numbered roles; each "role agent" competes to produce outputs distinct from the others, exploring complementary regions of the high-quality output space — no hand-crafted personas or architecture changes needed.
- A prompt-adaptive quality gating mechanism calibrates a reference quality threshold so diversity rewards only go to responses that meet it, preventing reward hacking.
- Evaluated on a benchmark suite spanning seven general capability tasks and four domain-specific diversity tasks.
More from Research
- Epoch AI audits 15 AI benchmarks: 9 flawed, only 4 verified — Jsevillamol · 2026-09-19
- Schmidhuber: his 1991 fast weights work was the first Transformer variant, 30 years before GPT — SchmidhuberAI · 2026-09-19
- Berkeley's PixelRAG hits 10k stars: screenshots beat text parsing for RAG retrieval — tom_doerr · 2026-09-19
- Terence Tao's SAIR nonprofit unveils Open Math Model initiative ahead of schedule — JosephJacks_ · 2026-09-19
- Talk on scaling fMRI foundation models: CortexMAE and Brainmarks — humanscotti · 2026-09-19
- SELF-INDEX: a framework for retrieval indexes that self-evolve without humans — Sangam Lee · 2026-09-19