LoRABuilding with AI

Low-rank adaptation

5 min readintermediateUpdated 28 Sept 2026
1 · In one line

LoRA adapts a large model to a new task by freezing its weights and training two thin matrices whose product is added to them.

1 · What it is

LoRA stands for low-rank adaptation. Researchers at Microsoft introduced it as a cheaper form of fine-tuning, which means further training an already trained model for one job. A model’s weights are the numbers it learned in training, stored in grids called matrices. LoRA does not edit those numbers. It leaves the original weight matrix frozen and learns the change as two smaller matrices, which it adds on top. The paper tested this on GPT-3, a Transformer. That is a model design built around attention: layers that let each word in a text draw on the other words. LoRA was applied only to the attention weights, and the rest stayed frozen.

Here is what “low rank” means. The original matrix W0 has d rows and k columns. LoRA adds two thin matrices, B and A. B is tall: d by r. A is wide: r by k. The number r is called the rank, and it is much smaller than d or k. Their product BA has the same shape as W0, but every column of BA is built by mixing the same r columns of B. So the whole change is built from just r basic patterns. The update is also scaled by α/r, where α is a fixed number. The authors note that this scaling means fewer other settings need re-tuning when r changes.

Why should a few patterns be enough? The authors guessed that the change a task needs has a low rank. Earlier work showed that pretrained models can be tuned using far fewer numbers than they contain. Training only 200 numbers, spread across the whole model through a fixed random mapping, got the RoBERTa language model to 90% of its fully tuned score on one benchmark, a standard test set.

Attention builds each word’s query (what it is looking for) and value (what it passes on) with learned matrices called Wq and Wv. Take one query matrix in GPT-3 175B. It is 12,288 by 12,288, about 151 million numbers. With rank 4, B and A together hold 98,304, roughly 1,500 times fewer. Counting every Wq and Wv matrix in all 96 layers, that comes to about 18 million trainable numbers. The saved task file shrank from 350 GB to about 35 MB, and training memory fell from 1.2 TB to 350 GB.

The best LoRA settings matched or beat full fine-tuning on all three GPT-3 benchmarks in the paper. WikiSQL, one of them, turns plain questions into SQL database queries. Even rank 1 on the query and value matrices was enough there.

The Hugging Face PEFT (parameter-efficient fine-tuning) library supports LoRA. In a PEFT documentation example, LoRA at rank 16 on the query and value matrices, plus a trained classifier head, left 667,493 of 86,543,818 numbers trainable, about 0.77%.

2 · Why it exists

Fully fine-tuning a big model retrains and copies every weight for every task.

Full-size copiesFull fine-tuning updates every parameter (one of the model's learned numbers), so each task gets a new model as large as the original. For GPT-3 that means 175 billion parameters per copy.
Training memoryFor every number it trains, training also stores a gradient (which way to nudge that number) and optimizer state (the update rule's own notes). In the LoRA paper, full fine-tuning of GPT-3 175B used 1.2 TB of GPU memory.
Older shortcuts costEarlier add-on methods either made the tuned model slower to answer or took up part of the text it can read at once. They often fell short of full fine-tuning.
3 · How it works

Follow one attention matrix in GPT-3 175B through training.

Only the thin pair learns. Its product has the full 12,288 × 12,288 shape, but rank at most 4.
  1. 1 · freezeKeep the pretrained weight matrix W0 fixed, so training never changes its numbers.
  2. 2 · attachAdd two thin matrices, A and B, whose shared inner size r (the rank) is tiny; B starts at zero, so at first the model is unchanged.
  3. 3 · trainThe input goes through both paths, the update's output is scaled by α/r, the two outputs are added, and training adjusts only A and B.
  4. 4 · mergeAfter training, add B times A into W0 once so the model answers as fast as before; to switch tasks, subtract it and add another pair.

LoRA rests on a bet: the change a task needs has low rank, even though the weight matrices themselves have full rank.

4 · Where it's used
WhoWhat they askWhat it works with
Support team“Can we adapt our assistant to a new task without retraining the whole model?”A small LoRA pair on the attention layers of a shared base model
Platform team“Can one base model serve several tuned tasks?”One frozen base with small per-task LoRA files swapped in
Image team“Can the image model be adapted with LoRA too?”A LoRA file loaded on top of the unchanged base model
5 · What it solves, and what it doesn't
solves
  • Compared with fully fine-tuning GPT-3 175B, it cuts trainable parameters by up to 10,000 times and GPU memory by 3 times.
  • One frozen base model can serve many tasks by swapping small LoRA files.
  • Merged weights add no extra delay when answering, unlike the older adapter method, which inserts extra layers into every Transformer block.
  • In a Llama-2-7B study, LoRA kept more of the base model's skills outside the target domain than full fine-tuning did.
doesn't solve
  • You still need the full base model to use a saved LoRA file.
  • On code and maths, the same study found LoRA in standard low-rank settings learned substantially less than full fine-tuning.
  • A small rank may not be enough when the task is far from pretraining, such as a different language.
  • Once a pair is merged into the weights, one batch of requests cannot easily mix tasks that need different LoRA pairs.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperLoRA: Low-Rank Adaptation of Large Language Models, Hu et al., Microsoft · read 28 Sept 2026
  2. repomicrosoft/LoRA: Code for loralib, an implementation of LoRA, Microsoft · read 28 Sept 2026
  3. docsLoRA (PEFT conceptual guide), Hugging Face · read 28 Sept 2026
  4. docsLoRA (PEFT reference), Hugging Face · read 28 Sept 2026
  5. paperIntrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning, Aghajanyan, Zettlemoyer and Gupta · read 28 Sept 2026
  6. paperLoRA Learns Less and Forgets Less, Biderman et al., TMLR 2024 · read 28 Sept 2026
  7. paperAttention Is All You Need, Vaswani et al., Google · read 28 Sept 2026