Training Agents 3: Reinforcement Learning
Hugging Face is teaching GRPO reinforcement learning for agent training in a live stream series
“it's pretty straightforward to learn from the available options and you can apply it on most use cases”
Hugging Face is running a multi-session live stream series on training AI agents, with this third session covering Group Relative Policy Optimization (GRPO), a reinforcement learning algorithm for updating model weights based on agent actions. The series progresses from supervised fine-tuning through distillation to RL, with a fourth session planned on building RL environments. This is practitioner-level educational content rather than a major product or research announcement.