DPO (Direct Preference Optimization)
DPO learns preferences from pairs of a better and a worse answer; it needs trustworthy preference data and adjusts weights.
When you will meet it
One representative preference-learning method; know it with RLHF to read most model cards.