To evaluate a fine-tuned LLM properly, the first decision is what to compare it with. The tempting comparison is the base model it started from, because that makes the training look good. The comparison that matters is the model you run in production today, on your task, with the prompt each one would actually get. This post covers that baseline, the held-out set, the metrics, the behaviours you did not tune, contamination, and the rollout decision.
Evaluate a fine-tuned LLM against the model you run today
A fine-tuned model is a replacement, so the question is whether the replacement is better than the incumbent. The incumbent is usually a larger hosted model with a carefully built prompt. The tuned model may run with a shorter prompt, because the examples moved into the weights. Give each side its own production prompt rather than a shared one; the comparison is between two systems, not two sets of weights.
Keep the base-versus-tuned comparison too, but for a different purpose. It tells you whether training did anything, which is useful for debugging the training run. It does not tell you whether to ship.
Run both on identical inputs, with the decoding settings you intend to use in production, and more than once where the output is sampled. Output varies run to run; a single pass through the test set can show a difference that disappears on the second pass.
Hold out a set before training starts
The held-out set is the portion of your data the model never sees, during training or during the tuning of hyperparameters. The order matters: split first, then train. A set carved out after the fact has usually been looked at, and a set you looked at while iterating is a development set, not a test set.
Split by something structural rather than by random row. Records from the same document, customer, repository or week are near-duplicates of each other; a random split puts one copy on each side and the score inflates. Runix Data splits evaluation from training data by source for this reason, and the same rule applies to data you build yourself.
Size it for the decision. A proportion measured over a few dozen items moves around enough to hide a real change; the same arithmetic that applies to refusal rates on live traffic applies here. Report a range, not a point.
Task-specific metrics versus generic benchmarks
Measure the task with the metric the task already has:
- Classification and routing: accuracy with a per-class breakdown, because an average hides the class you care about.
- Extraction: field-level exact match or overlap, and the share of outputs that parse against the schema at all.
- Code: whether the target tests pass with the change and the rest of the suite still passes.
- Generation without a gold answer: pairwise human judgement on a sample, with a model as judge only as a pre-filter and checked against the humans.
Generic benchmarks measure something else: broad capability on public question sets. Harnesses such as lm-evaluation-harness make them cheap to run, and they are worth running, but as a canary for general drift rather than as a target. A model tuned for your extraction task that drops a little on a public benchmark may be fine; one that drops a lot has probably forgotten something.
Check what you did not tune
Fine-tuning on a narrow set changes more than the narrow set. The behaviours most often damaged are the ones nobody wrote a test for: instruction following on other tasks, output in other languages, long-context handling, JSON for other schemas, and when the model declines.
Keep a small regression suite of exactly those behaviours, drawn from real traffic across routes you did not tune, and run it on every candidate. This matters doubly after preference tuning, which shifts refusals and format by design and can shift them further than intended. Decide the tolerated regression before you see the numbers, or the numbers will decide it for you.
Contamination: when the test set leaked into training
Contamination is any path by which test items reached the training process. The obvious one is a duplicate row. The less obvious ones are near-duplicates separated by random sampling, evaluation prompts reused to generate preference pairs or synthetic examples, and public benchmark items that were already in the base model's pretraining data.
Check for the first three with exact and near-duplicate matching between the test set and everything that touched training, and keep a record of which prompts were used for what. The last one you cannot fix, which is another argument for weighting your own held-out set above public scores.
When to roll out
- The tuned model beats the current one on the task metric by more than the run-to-run spread.
- The regression suite is inside the tolerance you set beforehand.
- The variant you will actually serve is the one you evaluated. A quantised build is a different model; evaluate it separately, and serve it only if quality holds. Merged adapters deserve the same re-check.
- Cost and latency were measured on the serving setup, as described in single-tenant model serving, not estimated from the training run.
- A slice of live traffic goes to the new model first, with the live signals watched and the previous model one configuration change away. A gateway such as Runix Router, where routing is set per key, can hold that pin so the switch is not a deploy.
Runix Models evaluates every tuned model on a task-specific held-out set, before and after, against the model you run today, and serves quantised variants only where that evaluation shows quality holds. It is in early access, by engagement.
Questions this raises
Should the fine-tuned model be compared with the base model or with the model in production?
Both, for different reasons. Base versus tuned tells you whether training worked; tuned versus the model you run today tells you whether to ship.
Are public benchmarks useful for a task-specific fine-tune?
As a canary for general capability loss, yes. As the target, no: they measure broad ability on public questions, not your task, and the base model may have seen them in pretraining.
What counts as contamination?
Any route by which test items reached training: duplicate rows, near-duplicates split at random, evaluation prompts reused to generate training or preference data, or public benchmark items already in the base model's pretraining.