GRPO
Software / App
A popular framework for post-training LLMs, with similarities in advantage calculation to RLP's reward mechanism.
Mentioned in 2 videos
Save the 2 videos on GRPO to your own pod.
Sign up free to keep building your knowledge base on GRPO as more episodes are added.
Videos Mentioning GRPO

Stanford CS25: Transformers United V6 I From Next-Token Prediction to Next-Generation Intelligence
Stanford Online
A popular framework for post-training LLMs, with similarities in advantage calculation to RLP's reward mechanism.

Stanford CS329A Self-Improving AI Agents | Part 6 | Train Time Scaling/Scaling RL
Stanford Online
A proposed technique in DeepSeek Math that reduces memory requirements in RL by using a generalized advantage estimation instead of a critic.