COLM 2026: metacognitive-reward RL, rubric self-distillation, and evolving research agents

armancohan · x · 2026-10-06

Three papers at COLM 2026:

The author invites discussions on RL, continual adaptation/learning, and training scientific agents.

Related event: Researcher Teases Three Papers at COLM 2026 Including Metacognitive Reward RL(2 posts)→

Original post →

More from Research

Research channel →