COLM 2026: metacognitive-reward RL, rubric self-distillation, and evolving research agents
armancohan · x · 2026-10-06
Three papers at COLM 2026:
- 'RL with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs' (Thursday 4:40pm, @pybeebee): trains LLMs to express uncertainty more faithfully via metacognitive rewards.
- 'Rethinking Reward Supervision: Rubric-Conditioned Self-Distillation' (Wednesday 11am): rethinking reward supervision with rubric-conditioned self-distillation.
- 'REVERE: Reflective Evolving Research Engineer' (Wednesday 11am): a reflective, evolving research agent.
The author invites discussions on RL, continual adaptation/learning, and training scientific agents.
More from Research
- Huawei Noah's Tail-Influence Sampling Cuts CVaR Policy Evaluation MSE by Up to 76% — huawei-noah · 2026-10-06
- Google's KeyRec Achieves Best Long-Video VLM Results With Just 10% of Visual Token Budget — google · 2026-10-06
- 4DCodeBench Shows Frontier Models Reconstruct Static Scenes but Fail at Dynamics — 4DCodeBench · 2026-10-06
- OmniTaskonomy: Year-long study shows generation training can improve understanding tasks — WeijiaShi2 · 2026-10-06
- Newton proved the product rule without limits, using a discrete symmetric-difference trick — ctjlewis · 2026-10-06
- DeepMind's AI designs enzymes from scratch: 99x drug building block yield, plastic-eating at 90°C — 141_1337 · 2026-10-06