New research shows user feedback signals like "that's wrong" can be leveraged for training
LChoshen · x · 2026-09-03
A new piece of work challenges the common view that feedback appearing in user responses ("that's wrong", "it worked!") is too noisy to leverage effectively. The researchers demonstrate these in-the-wild user signals can actually be used for training. Co-authored with LChoshen and Omri Abend; paper link in the original post.
More from Research
- LightOn ships NeoMME: natively multimodal encoders at 260M/800M with no vision tower — antoine_chaffin · 2026-09-03
- SOCO benchmark debuts at ECCV 2026: 1M+ pairs probe how vision models grasp object structure — HirokatuKataoka · 2026-09-03
- Two-Author Model Tech Report Praised as Dense: Pretraining to Downstream — antoine_chaffin · 2026-09-03
- Cohere Labs releases ATE dataset with ~700K tools from public MCP servers — Cohere_Labs · 2026-09-03
- IFM open-sources K2 Horizon models from 0.9B to 375B with training code and data recipes — testingcatalog · 2026-09-03
- Recursive LM author clarifies: what experiments prove vs. what the work is about — CShorten30 · 2026-09-03