NLP & Language · advanced · concept 98 of 176
DPO & Preference Optimization
Direct Preference Optimization tunes a model on pairs of preferred and rejected answers with a simple classification-style loss, no reward model and no RL loop. It made preference tuning accessible to anyone with a GPU, and most open models now align with DPO or one of its descendants.
Key terms
Preference pairsImplicit rewardReference modelDPO vs RLHFChosen and rejected
Where you meet it in the real world
Alignment of most open-weights chat models, style tuning, harmlessness training on a budget
Videos
▶ Stanford CS234 I Guest Lecture on DPO: Rafael Rafailov, Archit Sharma, Eric Mitchell I Lecture 9 ↗
Stanford Online · YouTube
Guides and articles
DPO Trainer · Hugging Face ↗
Hugging Face
Courses, papers, and more