DPO fine-tuning trains a model on pairs of answers, one preferred and one not, and moves it towards the preferred kind. It is the simplest way to teach behaviour that is easier to judge than to describe. This post covers what the data looks like, what the method actually optimises, how it differs from reinforcement learning with a reward model, and where it helps and where it cannot.
What preference data looks like
A preference record has three parts: a prompt, a chosen response, and a rejected response. Nothing says the chosen one is good in absolute terms, only that it is better than the other one for that prompt. The signal is a comparison, not a label.
That is the point. Writing a gold answer for "respond in our house tone" or "decline this kind of request politely" is hard; picking the better of two drafts is something a reviewer can do quickly and fairly consistently. Preference data captures judgement without requiring anyone to articulate the rule.
What DPO fine-tuning optimises
Direct Preference Optimization, from the paper by Rafailov and colleagues, starts from the same objective as reinforcement learning from human feedback: maximise a reward learned from preferences while keeping the model close to a reference model, with a coefficient, usually written beta, that sets how much drift is allowed.
The paper's observation is that this constrained objective has a closed-form solution, and that the reward implied by it can be written in terms of the model and the reference alone. Substituting that into the standard model of pairwise preferences turns the whole procedure into one classification loss over pairs:
loss = -log sigmoid(
beta * (log p(chosen) - log p_ref(chosen))
- beta * (log p(rejected) - log p_ref(rejected))
)
In words: raise the log-probability of the chosen response relative to the reference, lower it for the rejected response, and score the gap. The reference model is a frozen copy of the model you started from, usually the supervised fine-tuned model. It stays in memory, or its log-probabilities are precomputed, for the whole run.
How it differs from RLHF with a reward model
The earlier recipe, described in the InstructGPT paper, has three stages: supervised fine-tuning, training a separate reward model on preference pairs, then optimising the policy against that reward with a reinforcement learning algorithm while sampling new responses throughout. DPO keeps the first stage and collapses the other two.
- No reward model to train, host or debug. The policy is its own reward model, which is the claim in the paper's title.
- No sampling loop. DPO trains on fixed pairs like supervised learning; the RL recipe generates responses during training and scores them.
- Fewer moving parts. The main choice is
beta. The RL recipe has a reward model, a value model and an optimiser loop, each with its own instability.
What DPO gives up is on-policy exploration. The RL loop can discover a good response the data never contained; DPO never generates a candidate during training, so it learns only from the pairs it is given. In practice this makes the quality and coverage of the pairs the whole game, which is why the data section below is the longest. Implementations are available in the TRL library, and DPO combines with LoRA or QLoRA like any other fine-tuning step.
Where preference tuning helps
- Tone and register. Terse versus explanatory, formal versus plain: styles the model can already produce and only needs to prefer.
- Refusals. When to decline, and how. Pairs of a brusque refusal and a helpful one, or an over-refusal and a reasonable answer, teach the boundary better than a paragraph of policy in the prompt.
- Format adherence. Stopping after the answer, not restating the question, keeping to the schema. A chosen response that follows the format and a rejected one that nearly does is a precise signal.
- Choosing among valid answers. Tasks where several responses are correct and one is better, which supervised fine-tuning on a single gold answer cannot express.
Where it does not help
DPO fine-tuning does not add knowledge. It changes which of the responses the model could already produce it prefers; a fact the model does not have is not among them. For knowledge and freshness the tool is retrieval, as set out in fine-tuning versus RAG versus prompting.
It does not fix a capability gap either. If the supervised model cannot do the task at all, preferring better attempts over worse ones leaves you with a model that fails more politely. Do supervised fine-tuning first, confirm the model can do the task, and then use preferences to shape how it does it.
And it does not replace evaluation. The training loss going down tells you the model has separated your pairs; it does not tell you the behaviour transferred to new prompts, or that nothing else moved. Refusal rate and parse rate are exactly the signals preference tuning is most likely to shift, in either direction.
Collecting preference data without fooling yourself
- Sample pairs from the model you are tuning. Pairs copied from another model's outputs teach it to prefer responses it may never produce. Generate candidates from your supervised model and have those judged.
- Watch length. Reviewers tend to prefer the longer answer, and the model learns verbosity instead of quality. Compare the chosen and rejected length distributions before training.
- Change one thing per pair. A pair that differs in tone, length and correctness at once teaches a blend. Pairs that differ mainly in the behaviour you want are the useful ones.
- Measure rater agreement. Have several reviewers judge a shared subset. Low agreement means the pairs encode noise, and the model will learn the noise.
- Treat synthetic judges with suspicion. A model used as a judge brings its own preferences for style and length. Spot-check its decisions against humans before trusting it at volume.
- Keep the evaluation prompts out. If held-out prompts are used to generate pairs, the evaluation is contaminated before training starts. Runix Data keeps evaluation data split from training data by source for this reason.
Runix Models includes preference tuning with DPO alongside supervised fine-tuning, LoRA and QLoRA, applied to open-weight families and evaluated on a task-specific held-out set against the model you run today. It is in early access, by engagement; the data you provide trains and evaluates the model you commission, and nothing else.
Questions this raises
Does DPO need a reward model?
No. DPO derives the reward from the model and a frozen reference copy, so the only trained artefact is the policy itself. The reference model, or its log-probabilities, does need to be available during training.
Should DPO come before or after supervised fine-tuning?
After. Supervised fine-tuning gets the model to do the task; DPO shapes how it does it. Applied to a model that cannot do the task, preference tuning produces a model that fails more gracefully.
How many preference pairs are needed?
There is no fixed number; it depends on how consistent the pairs are and how narrow the behaviour is. Rater agreement and a clean held-out evaluation matter more than volume.