Glossary · Math & training

DPO (Direct Preference Optimization)

A preference-optimization objective that trains a policy directly from preferred and rejected response pairs relative to a reference policy. It avoids running an explicit reward model and reinforcement-learning loop during this stage.

Common confusion

DPO still depends on the quality and coverage of preference data and does not eliminate evaluation or alignment risk.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.