Yoav Goldberg: meaningless RL loss tweaks at COLM are no better than prompt engineering

yoavgo · x · 2026-10-11

NLP researcher Yoav Goldberg pushes back on a pervasive COLM talking point — "last year was boring, we only did prompt tweaking, now we're doing science again." He argues the supposed science is mostly meaningless tweaks to RL loss terms, no better than prompt engineering, which he calls perfectly fine work that people dismiss simply because it isn't math-y.

Related event: Yoav Goldberg Slams COLM Narrative: RL Loss Fine-Tuning Is Not Superior Science(2 posts)→

Original post →

More from Research

Research channel →