At COLM, researchers clash: RL-loss tweaks are no more 'science' than prompt tuning

BlancheMinerva · x · 2026-10-11

Yoav Goldberg cites a pervasive COLM sentiment—"last year was boring, we only did prompt tweaking, now we're doing science again"—and pushes back: that science is mostly meaningless tweaks to an RL loss term, no better than prompt tuning. tallinzen adds that this reveals a mismatch between researchers' training/identity and what 90% of impactful LLM work actually is: data, evals, policy, applications.

Related event: Yoav Goldberg Slams COLM Hype: RL Loss Tuning Is No More Scientific Than Prompts(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →