New agent-training livestream will teach GRPO with verifiers, reward hacking, and code-alongs
ben_burtenshaw · x · 2026-07-22
A new live-stream session on training agents will focus on GRPO and practical reinforcement-learning workflows.
- Earlier sessions covered imitation learning; this one moves to outcome-based training.
- The method samples multiple completions, scores them with a reward function, and uses the relative scores as the learning signal.
- The speaker emphasizes that GRPO can work without a reward model or critic, relying instead on simple verifiers and reward functions.
- The session will also cover common failure modes such as reward hacking, collapse, and length gaming, with code-along tutorials and reproducible experiments.
Related event: Hugging Face Live Stream Covers GRPO for Agent Reinforcement Learning(2 posts)→
More from coding & agent
- Agent harness providers are asked to support in-place tool updates without restarts — GabGarrett · 2026-07-22
- Bolt adds stacked team skills that can all trigger from one prompt — damianplayer · 2026-07-22
- A new RL finding says the training harness may induce generalization — inductionheads · 2026-07-22
- A Redditor Built a Local LLM Driven by Single-Photon Quantum Noise — Reddactor · 2026-07-22
- Coding agents are making stagnant enterprise software feel alive again — yunta_tsai · 2026-07-22
- Lakuna Router: Making AI Models 'Audition' for Real-World Workflows — Due_Hovercraft6497 · 2026-07-22