Yoav Goldberg: meaningless RL loss tweaks at COLM are no better than prompt engineering
yoavgo · x · 2026-10-11
NLP researcher Yoav Goldberg pushes back on a pervasive COLM talking point — "last year was boring, we only did prompt tweaking, now we're doing science again." He argues the supposed science is mostly meaningless tweaks to RL loss terms, no better than prompt engineering, which he calls perfectly fine work that people dismiss simply because it isn't math-y.
More from Research
- New arXiv paper examines scaling and emergent abstractions in byte-level language models — yogthos · 2026-10-11
- NVIDIA's GATOR turns casual photos into simulation-ready 3D objects with agentic refinement — AjayMandlekar · 2026-10-11
- Cowcraft MCP and WoWBench go live, testing LLM agents inside World of Warcraft — djcows · 2026-10-11
- Wuji opensources mjlab PPO stack for dexterous hand: in-hand reorientation and pen spinning — rohanpaul_ai · 2026-10-11
- Terence Tao's 27-slide deck: math is entering an era of proof abundance — ns123abc · 2026-10-11
- MemoType: type-based memory routing lifts agent recall by up to 16.18% — dair_ai · 2026-10-11