Fine-tuning vs RAG vs prompting: a decision guide by failure mode

Fine-tuning vs RAG vs prompting is usually framed as a choice between three ways of making a model better. It is more useful to treat them as three fixes for three different failures. A prompt changes what the model is asked, retrieval changes what it is shown, and fine-tuning changes what it is. Pick by diagnosing the failure, not by preference.

Three tools, three different problems

Prompting is instruction at inference time: the system message, the worked examples, the output format you ask for. It costs nothing to change and nothing to deploy, and every request pays for it again in input tokens.

Retrieval-augmented generation (RAG) fetches relevant text from a store you control and places it in the context before the model answers. The weights do not change; the information in front of the model does. The name comes from the original retrieval-augmented generation paper, which trained retriever and generator together; the production pattern keeps the weights fixed.

Fine-tuning continues training the model on examples of the task done well, so the behaviour lives in the weights rather than in the prompt. It is the only one of the three that produces a new model you have to evaluate, version and serve.

When prompting is enough

Start here, every time. A large share of "the model cannot do this" turns out to be "the model was not told clearly". Write the instruction you would give a careful new colleague, add a few worked examples, and state the output format exactly.

Prompting is enough when the model already has the knowledge and the skill, and the gap is in the asking. The signs: it gets the task right sometimes, gets it right when you rephrase, or gets it right after one correction in a longer conversation.

Its limits are cost and ceiling. A long prompt full of examples is paid for on every call, although prompt caching blunts that when the prefix is stable. And there is a ceiling: if the model does not know the fact, or cannot hold a style across thousands of requests, more instructions do not help.

When retrieval is the fix

Retrieval fixes knowledge problems. If the right answer depends on your documents, your product's current state, or anything that changed after the model's training data was collected, tuning puts it in the weights only unreliably, and only as of the training snapshot. The model has to be shown the text.

Retrieval is also the fix for freshness. A fine-tuned model is a snapshot; a retrieval index is updated by writing to it. If the facts change weekly, re-training weekly is the wrong shape of solution.

What retrieval does not fix: tone, format and reasoning habits. Putting the right passage in the context does not stop a model drifting into chatty prose. Retrieval also brings its own failure, the retriever returning the wrong passage, which looks exactly like the model being wrong until you log what was retrieved next to what was answered.

When fine-tuning is the right tool

Fine-tuning fixes behaviour that is consistent across requests and hard to specify in words. Four cases recur:

What fine-tuning does not fix: knowledge. Training on your documents teaches a model the shape of your documents far more reliably than their contents, and it does not keep up when the contents change. Teams that fine-tune for facts usually add retrieval later anyway.

Fine-tuning also carries a price that is not on the GPU bill: a dataset to build, a held-out set to protect, an evaluation to run, and a model to serve and keep serving. The methods differ in what they cost; LoRA, QLoRA and full fine-tuning covers what each one changes.

The cheapest fine-tuning vs RAG vs prompting experiment

Run the experiments in cost order. Each one produces the input for the next.

  1. Build a test set before touching anything. A few hundred real inputs with the answer you want for each. Without it, every later comparison is an anecdote.
  2. Prompt the strongest model you can run. If it passes, you have a ceiling and a baseline. If it fails, read the failures rather than counting them.
  3. Add retrieval if the failures are missing facts. Log what was retrieved alongside what was answered, so retriever errors and model errors can be told apart.
  4. Fine-tune a smaller model if the failures are behavioural. Compare it with the prompted model on the same test set, not with the base model it started from.

A frequent outcome is that the prompted large model is the right answer and the project ends at step two. That is a good result, not a failed one. The expensive mistake is starting at step four to find out.

Reading the failure mode

When a system is already live and underperforming, the failure usually names the tool it needs.

Mixed signals usually mean two gaps, and the tools stack: a tuned small model with retrieval in front of it is a common end state. Add them in the cost order above, and let an evaluation against the model you run today show that each step earned its place.

Runix Models is the fine-tuning branch of this guide: supervised fine-tuning, LoRA and QLoRA, and DPO on open-weight families, evaluated on a task-specific held-out set against the model you run today, with training data from your own team or from Runix Data. It is in early access, by engagement, scoped and quoted before any training starts, and the scoping begins with the question this post asks: whether a tuned model is the right tool for the failure in front of you, or whether a hosted model through Runix Router already is.

Questions this raises

Can fine-tuning and RAG be combined?

Yes, and it is a common end state: a tuned small model handles format and style while retrieval supplies the facts. Add them in cost order and measure each step against the same test set.

Does fine-tuning teach a model new facts?

Unreliably. Fine-tuning on documents teaches their style and structure far more reliably than their contents, and the contents go stale. Retrieval is the tool for knowledge and freshness.

How large should the test set be before choosing?

Large enough that the difference you care about is bigger than the run-to-run noise, which for most tasks means hundreds of real inputs rather than a dozen. Build it before the first experiment.

Related to this post: Runix Models. Tell us what you are building and we reply within one business day.