Agentic Research

DPO (Direct Preference Optimization)

DPO learns preferences from pairs of a better and a worse answer; it needs trustworthy preference data and adjusts weights.

When you will meet it

One representative preference-learning method; know it with RLHF to read most model cards.

Related terms

Next