Parameter-efficient fine-tuning
PEFT adapts a pretrained model by training a small set of added or selected parameters while keeping most base weights frozen.
Parameter-efficient fine-tuning is a group of methods, not one algorithm. Adapters, LoRA, prefix tuning and prompt tuning all keep the pretrained weights frozen and train only a small set of extra numbers. They differ in where those numbers sit. Prompt and prefix tuning add learned vectors in front of what the model reads. Adapters and LoRA add small matrices alongside the model’s own weights.
LoRA writes the change to a weight matrix as two thin matrices multiplied together. During training the original matrix stays exactly as it was; only the two thin matrices change. At serving time, their product can be added into the frozen weight, so there is no extra step to slow inference.
What you gain is a smaller memory bill during training, and usually a smaller compute bill too. For language models, Hugging Face suggests a ladder. Start by checking whether plain prompting, such as a few worked examples, already does the job. If it does, a learned prompt can do that job with fewer tokens. If a learned prompt isn’t enough, step up to adapters or layer tuning. Those have more knobs to turn than a prompt, so their results usually land nearer to a model retrained in full. Whatever the method, test whether the tuned model still knows what the base model knew.
Full fine-tuning updates and stores a complete parameter set for every task, which becomes costly as base models grow.
Freeze the backbone and train a compact task-specific path.
- 1 · freezeLoad a pretrained model and keep its original weights fixed.
- 2 · attachClip small, learnable matrices onto the frozen model.
- 3 · trainRun training as usual, but let the optimizer change only the new pieces.
- 4 · deployFold the small update into the frozen weights before serving. Or keep it separate and swap adapters by task.
Parameter-efficient means fewer trainable parameters. It is not a promise of full fine-tuning's accuracy.
| Who | What they ask | What it works with |
|---|---|---|
| Model platform | “Can many customer adaptations share one base checkpoint?” | Adapter size and hot-swap latency |
| Training team | “Which modules and rank should LoRA target?” | Quality versus trainable parameter count |
| Evaluator | “Did adaptation erase base capabilities?” | Task gain and retention benchmarks |
- Frozen weights skip gradient math and optimizer bookkeeping, so LoRA training fits in far less GPU memory.
- Task files get tiny; LoRA shrank one GPT-3 task checkpoint roughly 10,000 times.
- Many tasks can share one frozen base model, each with its own small add-on.
- With less room to learn, a tuned model may fall short of what full fine-tuning would reach.
- The tuned model can still lose skills the base model had, so check what it kept.
- Most methods add some cost when the model runs; folding an adapter into the base weights removes part of it.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperLoRA: Low-Rank Adaptation of Large Language Models, Hu et al. · read 28 Sept 2026
- paperParameter-Efficient Transfer Learning for NLP, Houlsby et al. · read 28 Sept 2026
- paperThe Power of Scale for Parameter-Efficient Prompt Tuning, Lester et al. · read 28 Sept 2026
- paperPrefix-Tuning: Optimizing Continuous Prompts for Generation, Li and Liang · read 28 Sept 2026
- docsParameter efficient fine-tuning methods, Hugging Face · read 28 Sept 2026