Quantized low-rank adaptation
QLoRA fine-tunes a large language model by storing its frozen weights in 4 bits and training only a small add-on that sits beside them.
QLoRA, short for quantized low-rank adaptation, comes from a May 2023 paper by researchers at the University of Washington.
Start with quantization. It means storing numbers with fewer bits, so each one takes less memory. For example, you can turn 32-bit decimal numbers into 8-bit whole numbers. Think of rounding prices to the nearest dollar: you lose a little detail but save a lot of space.
Next comes LoRA. LoRA freezes a model’s original weights and learns only a small change to them. “Low-rank” describes the shape of that change. Instead of a full grid as big as a weight matrix W, the change is two thin matrices, L1 and L2. Multiply them together and add the result to W. A small number s scales it. QLoRA keeps the frozen W in 4 bits and trains only L1 and L2.
The 4-bit format is called NF4, or 4-bit NormalFloat. Trained weights usually form a bell curve centred on zero. NF4 spaces its levels so each one covers an equal share of that bell curve. Each block of 64 weights also needs its own scale number. With 32-bit scales, that adds 0.5 bits per weight. A trick called double quantization stores those scales in 8 bits instead. That cuts the extra cost to about 0.127 bits per weight and saves roughly 3 GB on a 65B model.
Here is where the memory goes. For a 7B LLaMA model, the 4-bit base took 5,048 MB. The adapters held about 0.2% as many numbers as the model and took only 26 MB. Since adapters are so cheap, the authors put them on every linear layer. They found this was needed to match full fine-tuning.
To test it, the authors trained a family of chatbots called Guanaco with QLoRA. They scored them on the Vicuna benchmark: 80 prompts, with GPT-4 judging each answer against ChatGPT’s. The largest Guanaco reached 99.3% of ChatGPT’s score after about a day on one GPU. Treat that number with care. GPT-4 tended to favour whichever answer it saw first, and the authors say tests like this are not a reliable way to rank chatbots. A 33B version fits on a 24 GB consumer GPU and trains in under half a day.
In Hugging Face code, QLoRA is a loading config with 4-bit loading, the nf4 type and double quantization switched on. A LoRA config then adds the adapters.
Fine-tuning a big model in 16-bit needs hundreds of gigabytes of GPU memory.
Follow one layer of the model through a training step.
- 1 · quantizeThe pretrained weights are frozen and stored as 4-bit NF4 codes, in blocks of 64 that each carry their own scale.
- 2 · attachTwo thin trainable matrices, L1 and L2, kept in 16-bit, are added beside every linear layer.
- 3 · computeWhen a layer runs, its 4-bit weights are unpacked to BF16 (16-bit BrainFloat, the number format used for the maths), multiplied with the input, and the adapter's output is added.
- 4 · updateGradients, the error signals that say how weights should change, pass back through the frozen weights, but only the adapter changes.
The Q is the frozen 4-bit copy of the model; the LoRA is the small part that actually learns.
| Who | What they ask | What it works with |
|---|---|---|
| Researcher with one GPU | “Can I adapt a 33B model on a single 24 GB card?” | A 4-bit base model with LoRA adapters |
| Chatbot builder | “Does a small, high-quality instruction dataset beat a much larger one?” | Instruction data such as OASST1 |
| ML engineer | “Will a 13B model fine-tune on a 16 GB T4?” | NF4 weights with nested quantization |
| Alignment team | “Can one base model carry separate adapters for separate jobs?” | One 4-bit base with several adapters |
- The paper fine-tuned a 65B model on one 48 GB GPU, where regular 16-bit fine-tuning needs over 780 GB.
- In the authors' tests, QLoRA replicated 16-bit full fine-tuning results, and matched 16-bit LoRA on the MMLU benchmark.
- Paged optimizers move the optimizer state (extra numbers the training process keeps for the weights it updates) to the computer's main memory when the GPU runs out of room.
- It is built into bitsandbytes, a quantization library for PyTorch, and into Hugging Face's model libraries.
- For the biggest models tried, 33B and 65B, the paper leaves open whether QLoRA is as good as full 16-bit fine-tuning.
- It does not allow pure 4-bit training; only extra parameters such as adapters can be trained.
- It saves memory, not maths. The multiplications still run in 16-bit, never in 4-bit.
- Only one recipe was studied, 4-bit base weights with LoRA. Whether 3-bit weights or other kinds of adapter work as well is untested.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperQLoRA: Efficient Finetuning of Quantized LLMs, Dettmers et al., University of Washington · read 28 Sept 2026
- paperLoRA: Low-Rank Adaptation of Large Language Models, Hu et al. · read 28 Sept 2026
- docsbitsandbytes, Hugging Face · read 28 Sept 2026
- docs4-bit quantization (bitsandbytes API reference), Hugging Face · read 28 Sept 2026
- docsQuantization (PEFT documentation), Hugging Face · read 28 Sept 2026
- docsBitsandbytes (Transformers documentation), Hugging Face · read 28 Sept 2026
- officialMaking LLMs even more accessible with bitsandbytes, 4-bit quantization and QLoRA, Hugging Face · read 28 Sept 2026
- repoartidoro/qlora, UW NLP · read 28 Sept 2026