Microsoft Research proposes LLM-as-a-Coach for non-verifiable tasks
donglixp · x · 2026-07-22
Microsoft Research’s LLM-as-a-Coach paper proposes a new way to train models on non-verifiable tasks.
- The authors argue that scalar rewards throw away rich feedback in open-ended tasks.
- Their method, Experiential Learning (EL), repurposes the feedback model from an LLM-as-a-Judge setup into an LLM-as-a-Coach.
- The coach turns per-response assessments into transferable experiential knowledge, which is then used to condition a teacher model and distilled into the policy through on-policy context distillation.
- Compared with standard rubric-based RL, EL reportedly preserves fine-grained preferences, provides denser supervision, generalizes better out of distribution, and reduces reward hacking.
- The paper says EL outperforms rubric-based RL across two policy families, using either self-feedback or proprietary-model feedback.
Related event: Microsoft Proposes LLM-as-a-Coach for Non-Verifiable Tasks(4 posts)→
More from Research
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11
- Fruit fly connectome LLM weights land on Hugging Face, transformers-compatible — ngxson · 2026-09-11
- Fruit fly brain as an LLM: connectome-driven language model demo goes live — ngxson · 2026-09-11
- Harry Collins: LLMs can't do frontier science because they can't invent new language — whoamisri · 2026-09-11
- The Waymo effect: how AI is quietly making research less collaborative — JohnHammersley · 2026-09-11
- Causal-only attention for non-generative tasks is wasteful, argues HF engineer — antoine_chaffin · 2026-09-11