QLoRABuilding with AI

Quantized low-rank adaptation

5 min readintermediateUpdated 28 Sept 2026
1 · In one line

QLoRA fine-tunes a large language model by storing its frozen weights in 4 bits and training only a small add-on that sits beside them.

1 · What it is

QLoRA, short for quantized low-rank adaptation, comes from a May 2023 paper by researchers at the University of Washington.

Start with quantization. It means storing numbers with fewer bits, so each one takes less memory. For example, you can turn 32-bit decimal numbers into 8-bit whole numbers. Think of rounding prices to the nearest dollar: you lose a little detail but save a lot of space.

Next comes LoRA. LoRA freezes a model’s original weights and learns only a small change to them. “Low-rank” describes the shape of that change. Instead of a full grid as big as a weight matrix W, the change is two thin matrices, L1 and L2. Multiply them together and add the result to W. A small number s scales it. QLoRA keeps the frozen W in 4 bits and trains only L1 and L2.

The 4-bit format is called NF4, or 4-bit NormalFloat. Trained weights usually form a bell curve centred on zero. NF4 spaces its levels so each one covers an equal share of that bell curve. Each block of 64 weights also needs its own scale number. With 32-bit scales, that adds 0.5 bits per weight. A trick called double quantization stores those scales in 8 bits instead. That cuts the extra cost to about 0.127 bits per weight and saves roughly 3 GB on a 65B model.

Here is where the memory goes. For a 7B LLaMA model, the 4-bit base took 5,048 MB. The adapters held about 0.2% as many numbers as the model and took only 26 MB. Since adapters are so cheap, the authors put them on every linear layer. They found this was needed to match full fine-tuning.

To test it, the authors trained a family of chatbots called Guanaco with QLoRA. They scored them on the Vicuna benchmark: 80 prompts, with GPT-4 judging each answer against ChatGPT’s. The largest Guanaco reached 99.3% of ChatGPT’s score after about a day on one GPU. Treat that number with care. GPT-4 tended to favour whichever answer it saw first, and the authors say tests like this are not a reliable way to rank chatbots. A 33B version fits on a 24 GB consumer GPU and trains in under half a day.

In Hugging Face code, QLoRA is a loading config with 4-bit loading, the nf4 type and double quantization switched on. A LoRA config then adds the adapters.

2 · Why it exists

Fine-tuning a big model in 16-bit needs hundreds of gigabytes of GPU memory.

Huge memory billRegular 16-bit fine-tuning of a 65-billion-parameter LLaMA model needs over 780 GB of GPU memory.
Compression broke trainingEarlier quantization methods shrank models for running them, but broke down when used for training.
LoRA still stores 16-bitLoRA already freezes the base model and trains small added matrices. QLoRA goes further by storing that frozen base in 4-bit precision.
3 · How it works

Follow one layer of the model through a training step.

Only the adapter path learns. The frozen 4-bit weights are unpacked to 16-bit each time the layer runs.
  1. 1 · quantizeThe pretrained weights are frozen and stored as 4-bit NF4 codes, in blocks of 64 that each carry their own scale.
  2. 2 · attachTwo thin trainable matrices, L1 and L2, kept in 16-bit, are added beside every linear layer.
  3. 3 · computeWhen a layer runs, its 4-bit weights are unpacked to BF16 (16-bit BrainFloat, the number format used for the maths), multiplied with the input, and the adapter's output is added.
  4. 4 · updateGradients, the error signals that say how weights should change, pass back through the frozen weights, but only the adapter changes.

The Q is the frozen 4-bit copy of the model; the LoRA is the small part that actually learns.

4 · Where it's used
WhoWhat they askWhat it works with
Researcher with one GPU“Can I adapt a 33B model on a single 24 GB card?”A 4-bit base model with LoRA adapters
Chatbot builder“Does a small, high-quality instruction dataset beat a much larger one?”Instruction data such as OASST1
ML engineer“Will a 13B model fine-tune on a 16 GB T4?”NF4 weights with nested quantization
Alignment team“Can one base model carry separate adapters for separate jobs?”One 4-bit base with several adapters
5 · What it solves, and what it doesn't
solves
  • The paper fine-tuned a 65B model on one 48 GB GPU, where regular 16-bit fine-tuning needs over 780 GB.
  • In the authors' tests, QLoRA replicated 16-bit full fine-tuning results, and matched 16-bit LoRA on the MMLU benchmark.
  • Paged optimizers move the optimizer state (extra numbers the training process keeps for the weights it updates) to the computer's main memory when the GPU runs out of room.
  • It is built into bitsandbytes, a quantization library for PyTorch, and into Hugging Face's model libraries.
doesn't solve
  • For the biggest models tried, 33B and 65B, the paper leaves open whether QLoRA is as good as full 16-bit fine-tuning.
  • It does not allow pure 4-bit training; only extra parameters such as adapters can be trained.
  • It saves memory, not maths. The multiplications still run in 16-bit, never in 4-bit.
  • Only one recipe was studied, 4-bit base weights with LoRA. Whether 3-bit weights or other kinds of adapter work as well is untested.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperQLoRA: Efficient Finetuning of Quantized LLMs, Dettmers et al., University of Washington · read 28 Sept 2026
  2. paperLoRA: Low-Rank Adaptation of Large Language Models, Hu et al. · read 28 Sept 2026
  3. docsbitsandbytes, Hugging Face · read 28 Sept 2026
  4. docs4-bit quantization (bitsandbytes API reference), Hugging Face · read 28 Sept 2026
  5. docsQuantization (PEFT documentation), Hugging Face · read 28 Sept 2026
  6. docsBitsandbytes (Transformers documentation), Hugging Face · read 28 Sept 2026
  7. officialMaking LLMs even more accessible with bitsandbytes, 4-bit quantization and QLoRA, Hugging Face · read 28 Sept 2026
  8. repoartidoro/qlora, UW NLP · read 28 Sept 2026