CMU L3 Lab brings GradAlign RL data selection and sim2real papers to COLM 2026
wellecks · x · 2026-10-06
Three PhD students from CMU's L3 Lab (Sean Welleck's group) are presenting at COLM 2026:
- GradAlign: gradient-aligned data selection for LLM RL. Targeting the non-stationarity problem where RL training is highly sensitive to problem quality, it uses a small trusted validation set to prioritize training problems whose policy gradients align with validation gradients, forming an adaptive curriculum. It consistently beats baselines across unreliable rewards, distribution imbalance, and low-utility corpora. Code released at github.com/StigLidu/GradAlign
- Mind the Sim2Real Gap in User Simulation for Agentic Tasks: examines the sim2real gap when using LLM-based user simulators for multi-turn agent evaluation
More from Research
- Studies: humans deny AI consciousness even with identical behavior; AI vision misses illusions primates catch — MacrinePhD · 2026-10-06
- SFT then RL doesn't fix agent looping: 29% of runs hit turn cap vs 0% for RL alone — VikParuchuri · 2026-10-06
- RL Post-Training Eliminates Agent Tool-Call Loops: 92% Loop Rate Drops to 0 — VikParuchuri · 2026-10-06
- Watch, Infer, Coordinate: robots infer a partner's physical limits from watching teamwork, then coordinate zero-shot — mangahomanga · 2026-10-06
- Swapping AdamW States for FFT Cuts Fine-tuning VRAM by 50% Without Quantization — Spectra-Global · 2026-10-06
- Math lacks empirical tradition: Wolfskehl Prize drew 1,000 wrong Fermat proofs — RexDouglass · 2026-10-06