Encyclopedia · 176 concepts

NLP & Language · advanced · concept 98 of 176

DPO & Preference Optimization

Direct Preference Optimization tunes a model on pairs of preferred and rejected answers with a simple classification-style loss, no reward model and no RL loop. It made preference tuning accessible to anyone with a GPU, and most open models now align with DPO or one of its descendants.

Key terms

Preference pairsImplicit rewardReference modelDPO vs RLHFChosen and rejected

Where you meet it in the real world

Alignment of most open-weights chat models, style tuning, harmlessness training on a budget