Fine-tuning
Fine-tuning takes a model that is already trained and keeps training it on a smaller set of examples for one task.
Fine-tuning means taking a model that has already been trained and training it a little more, on a smaller dataset for one task. The first round of training is called pretraining. It starts from random weights. Fine-tuning starts from the weights the model already learned. That is why it needs far less computing power, data and time.
Here is the recipe from the Keras guide, with a model that sorts photos of cats and dogs. First you freeze the pretrained layers, which means their weights cannot change. You add a new final layer for your two classes and train only that. Then you let some or all of those layers change again. You keep training, but the learning rate is set very low, so the weights change only a little at a time. The Keras guide warns that big weight updates on a small dataset risk overfitting very quickly.
The same idea works for language. In the ULMFiT paper, a fine-tuned language model given only 100 labelled examples matched one trained from scratch on 100 times more data. There is a cost, though. In full fine-tuning every weight can change, so each tuned copy is as big as the original. Methods such as LoRA freeze the original weights and train small add-on pieces instead.
Training a big model from nothing for every new job is slow and needs lots of data.
Follow a cat-or-dog photo model as it is fine-tuned.
- 1 · loadLoad a model whose weights were already learned on a large dataset.
- 2 · freezeFreeze its layers, so their weights do not change during training.
- 3 · headAdd a new final layer for your task and train only that layer.
- 4 · unfreezeLet some or all base layers change again, then retrain while keeping the learning rate very low.
- 5 · watchKeep weight changes small, because big ones on a small dataset risk overfitting.
Fine-tuning starts from learned weights, not random ones, and changes them gently.
| Who | What they ask | What it works with |
|---|---|---|
| Photo app team | “Can a general image model sort cats from dogs” | A small set of labelled pet photos |
| Developer tools team | “Can the model get better at writing code” | A dataset of coding examples |
| Support team | “Can replies match our style more reliably” | Example prompts paired with approved replies |
- It lets a model reuse what it learned before instead of starting over.
- It needs far less compute, data and time than training from scratch.
- It can make a model more reliably produce the style and content you want.
- It can overfit quickly when the new dataset is small.
- Full fine-tuning of a huge model still needs a lot of hardware, and each copy is as large as the original.
- It does not prove itself; you need tests that compare it with the base model.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsFine-tuning, Hugging Face · read 28 Sept 2026
- docsTransfer learning & fine-tuning, Keras · read 28 Sept 2026
- docsTransfer learning and fine-tuning, TensorFlow · read 28 Sept 2026
- paperUniversal Language Model Fine-tuning for Text Classification, Howard and Ruder · read 28 Sept 2026
- paperLoRA: Low-Rank Adaptation of Large Language Models, Hu et al., Microsoft · read 28 Sept 2026
- docsSupervised fine-tuning, OpenAI · read 28 Sept 2026