DPO

Software / App

Direct Preference Optimization, a simpler RLHF algorithm that eliminates the need for a separate reward model and on-policy sampling. It works by taking gradient steps towards preferred responses and negative steps away from dispreferred ones.

Mentioned in 2 videos

Save the 2 videos on DPO to your own pod.

Sign up free to keep building your knowledge base on DPO as more episodes are added.

Get Started Free