LoRA vs QLoRA vs full fine-tuning: what each one changes in the weights

LoRA vs QLoRA is a question about memory first and quality second, and the honest answer depends on which one you are short of. Both are ways of fine-tuning a model without updating all of its weights. This post sets out what each one changes, what that does to the GPU bill, what it costs in quality, and when training every weight is still the right call.

What full fine-tuning changes

Full fine-tuning updates every parameter in the model. For each trainable parameter, training keeps the weight itself, its gradient, and the optimiser state; with the Adam family of optimisers that state is two extra values per parameter. Activations for the backward pass sit on top.

That is why the memory needed to train a model is a multiple of the memory needed to run it. For a small model on a large accelerator it is fine. For a mid-size model it is the reason teams reach for adapters before they reach for more hardware.

What LoRA changes instead

LoRA, from the paper LoRA: Low-Rank Adaptation of Large Language Models, freezes the pretrained weights and trains a small update beside them. For a weight matrix W, the update is written as the product of two thin matrices, B and A, with a shared inner dimension called the rank. The effective weight at inference is W + BA.

The rank is small compared with the matrix it adapts, so the number of trainable parameters falls by orders of magnitude. The paper reports that, against GPT-3 175B fine-tuned with Adam, LoRA cut trainable parameters by a factor of about ten thousand and GPU memory by about three times. Only the adapter needs gradients and optimiser state; the frozen base still has to be in memory, but at inference precision, and activations are still kept for the backward pass, which is why the memory saving is a few times rather than ten thousand.

Two details matter in practice. B starts at zero, so the model begins training as exactly the base model. And the adapter can be applied to some weight matrices and not others; the paper's experiments adapt the attention projections, and the choice of which matrices to adapt, together with the rank, is most of the tuning space.

# illustrative LoRA configuration with the PEFT library
from peft import LoraConfig, get_peft_model
from transformers import AutoModelForCausalLM

base_model = AutoModelForCausalLM.from_pretrained("your-base-model")
config = LoraConfig(
    r=16, lora_alpha=32, lora_dropout=0.05,
    target_modules=["q_proj", "v_proj"],
    task_type="CAUSAL_LM",
)
model = get_peft_model(base_model, config)
model.print_trainable_parameters()

What QLoRA adds, and what it costs

QLoRA, from the paper QLoRA: Efficient Finetuning of Quantized LLMs, keeps the LoRA idea and shrinks the one thing LoRA leaves large: the frozen base. The base weights are stored in a 4-bit format the authors call NormalFloat, with a second quantisation pass over the quantisation constants themselves, and gradients flow through the quantised base into ordinary LoRA adapters kept at higher precision. The paper reports fine-tuning a 65B parameter model on a single 48GB GPU this way.

The cost is compute. The 4-bit weights are dequantised on every forward and backward pass, so a QLoRA step does more work than a LoRA step on the same hardware, and training runs take longer. The paper also introduces paged optimiser state to survive memory spikes rather than crash on them.

LoRA vs QLoRA in memory, compute and quality

On quality, LoRA's low rank limits how far the weights can move. That is a feature when you want to keep the base model's general ability and a limitation when the task needs a lot that the base does not have. The LoRA paper reports results on par with full fine-tuning on its benchmarks; whether that holds for your task is an empirical question, which is why the held-out evaluation is not optional.

QLoRA's authors report matching 16-bit fine-tuning quality in their experiments. Treat that as a claim to verify on your data rather than a property: the quantised base is not the same model as the full-precision base, and the adapter learned to work with the quantised one.

Merge the adapter or serve it separately

An adapter can be merged into the base by computing W + BA once and saving the result. The merged model serves like any other, with no extra computation per token, which is the property the LoRA paper highlights. The price is one full model per task.

Kept separate, one base can carry many adapters, and some serving engines can switch adapters per request. That suits a fleet of narrow tasks on one base and costs a small amount of extra computation on the adapter path.

One caveat for QLoRA: merging an adapter trained against the 4-bit base into the full-precision weights changes what it was trained against. Evaluate the merged model again rather than assuming the training numbers carry over, and if you intend to serve a quantised variant, evaluate that variant too. Where the model ends up is covered in single-tenant model serving.

When full fine-tuning is still warranted

Full fine-tuning has its own failure, which is forgetting: a model trained hard on a narrow set can lose general ability it had before. Mixing in general data and checking behaviours you did not tune are part of the job. Whether you should be tuning at all, rather than prompting or retrieving, is the question the fine-tuning versus RAG versus prompting guide is for. The tooling most teams start from is the PEFT library.

Runix Models offers supervised fine-tuning, LoRA and QLoRA on open-weight families, with the method chosen per engagement and the result evaluated on a held-out set against the model you run today. It is in early access, by engagement, and quantised variants are served only where that evaluation shows quality holds.

Questions this raises

Is QLoRA lower quality than LoRA?

Not necessarily, but it is a different starting point: the adapter is trained against a 4-bit base. The QLoRA authors report matching 16-bit fine-tuning in their experiments; confirm it on your own held-out set.

Can a LoRA adapter be served without merging it?

Yes. Some serving engines load adapters next to a shared base and switch per request, at a small cost in computation per token. Merging gives a plain model with no adapter path.

When is full fine-tuning worth the memory?

When the change is large, such as a new language or domain, when you add tokens, when the model is small enough to fit anyway, or when adapters have plateaued on your evaluation.

Related to this post: Runix Models. Tell us what you are building and we reply within one business day.