Three LLM research papers coming at COLM 2026: metacognitive RL, rubric rewards, research agents
armancohan · x · 2026-10-06
The author previews three papers to be presented with collaborators at COLM 2026: RL with metacognitive rewards, opsd with rubric rewards, and evolving research agents, with details in the follow-up thread.
More from Research
- Cohere Labs at COLM: Chain-of-Thought Legibility Is Not Real Interpretability — Cohere_Labs · 2026-10-06
- NVIDIA Research Hiring Scientists, Interns for World Models and Diffusion LMs — ArashVahdat · 2026-10-06
- Gautam Kamath: math community's 'human understanding' push makes AI proofs worth revisiting — thegautamkamath · 2026-10-06
- Everyone uses AI for new math results; Kamath wants AI to simplify old proofs — thegautamkamath · 2026-10-06
- PlurPO: training LLMs to curb social sycophancy that discourages relationship repair — RishiBommasani · 2026-10-06
- First large-scale 3B/8B continuous diffusion LMs match pass@1 and beat pass@k vs masked dLMs — ArashVahdat · 2026-10-06