LLM-as-a-Coach turns judge feedback into transferable experiential knowledge
iScienceLuvr · x · 2026-07-21
The paper proposes Experiential Learning (EL), which turns the usual LLM-as-a-Judge setup into an LLM-as-a-Coach.
Instead of collapsing feedback into a scalar score, the coach distills its assessment of each on-policy response into transferable experiential knowledge. That knowledge conditions a teacher model and is then internalized by the policy through on-policy context distillation. The authors report that, compared with standard rubric-based RL, EL improves held-out and unseen open-ended tasks, generalizes better beyond the training distribution, and helps mitigate reward hacking.
The screenshot also shows the paper’s high-level claim: richer, higher-bandwidth feedback preserves fine-grained preferences among high-quality responses and works across two policy families.
Related event: Microsoft Proposes LLM-as-a-Coach for Non-Verifiable Tasks(4 posts)→
More from Research
- 3D ResNet Paper Crosses 3,000 Citations Eight Years After CVPR 2018 — HirokatuKataoka · 2026-09-11
- Jeff Heaton's Intro to the Math of Neural Networks eBook Is Free to Download — blaizedsouza · 2026-09-11
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11