Parameter-Efficient Fine-Tuning
PEFT is a free Hugging Face library that teaches a big AI model a new task by training a small set of extra weights instead of the whole model.
PEFT is a free, open-source library from Hugging Face. The name stands for parameter-efficient fine-tuning. Fine-tuning means taking a model that is already trained and teaching it a new job. PEFT does this without changing most of the model. It freezes the original weights, which you can think of as the numbers the model already knows. Then it trains a small set of new ones.
The savings can be large. In the library’s own quick example, about 524 thousand weights are trained out of about 1.2 billion. Only the new weights get saved, and this small add-on is called an adapter. An adapter is often a few megabytes. The full model can be hundreds of megabytes or more. Quality is reported to be comparable to retraining the whole model.
One method PEFT supports is LoRA, short for low-rank adaptation. LoRA describes the change to each chosen layer with two small grids of numbers instead of one big one. You can keep several adapters for different tasks on one base model and switch between them. You can also merge one into the model so it adds no extra delay when answering.
Retraining a whole large model for each new job causes two problems.
Follow one model as PEFT adapts it to a new task.
- 1 · loadYou load a pretrained base model, for example from the Transformers library.
- 2 · configureA config names the method, such as LoRA, and which layers to adapt.
- 3 · wrapget_peft_model freezes the targeted layers and adds new, small trainable weights to them.
- 4 · trainTraining changes only the new weights, a tiny share of the total.
- 5 · saveYou save just the adapter, a file of megabytes instead of gigabytes.
The base model stays the same; only the small adapter is new.
| Who | What they ask | What it works with |
|---|---|---|
| Student | “Can I fine-tune a language model on one graphics card at home?” | LoRA or QLoRA through PEFT on consumer hardware |
| App developer | “Can one base model serve several tasks without storing several full copies?” | Several small adapters loaded on one base model |
| Diffusers user | “Does PEFT work with the Diffusers library too?” | PEFT's built-in Diffusers integration |
- It trains only a small number of extra weights, cutting compute and storage costs.
- Its adapter files are usually megabytes, not gigabytes.
- One base model can hold several adapters and switch between them.
- It works with the Transformers, Diffusers and Accelerate libraries.
- Results are comparable to full fine-tuning, not guaranteed to be identical.
- It does not supply training data or decide what the model should learn.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsPEFT, Hugging Face · read 28 Sept 2026
- repohuggingface/peft: State-of-the-art Parameter-Efficient Fine-Tuning, Hugging Face · read 28 Sept 2026
- docsQuicktour, Hugging Face · read 28 Sept 2026
- docsAdapters, Hugging Face · read 28 Sept 2026
- paperLoRA: Low-Rank Adaptation of Large Language Models, arXiv · read 28 Sept 2026
- docsPEFT checkpoint format, Hugging Face · read 28 Sept 2026