Hugging Face Live Stream Covers GRPO for Agent Reinforcement Learning
Hugging Face's latest live stream focuses on training agents using GRPO, transitioning from imitation learning to outcome-based reinforcement learning. The session covers practical workflows, verifiers, and preventing reward hacking.
2026-07-22 ~ 2026-07-22 · 2 related posts
- Hugging Face Training Agents 3 explains GRPO for outcome-based agent training — Hugging Face · 2026-07-22
- New agent-training livestream will teach GRPO with verifiers, reward hacking, and code-alongs — ben_burtenshaw · 2026-07-22