Concepts

Fine-tuning

4 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Fine-tuning takes a model that is already trained and keeps training it on a smaller set of examples for one task.

1 · What it is

Fine-tuning means taking a model that has already been trained and training it a little more, on a smaller dataset for one task. The first round of training is called pretraining. It starts from random weights. Fine-tuning starts from the weights the model already learned. That is why it needs far less computing power, data and time.

Here is the recipe from the Keras guide, with a model that sorts photos of cats and dogs. First you freeze the pretrained layers, which means their weights cannot change. You add a new final layer for your two classes and train only that. Then you let some or all of those layers change again. You keep training, but the learning rate is set very low, so the weights change only a little at a time. The Keras guide warns that big weight updates on a small dataset risk overfitting very quickly.

The same idea works for language. In the ULMFiT paper, a fine-tuned language model given only 100 labelled examples matched one trained from scratch on 100 times more data. There is a cost, though. In full fine-tuning every weight can change, so each tuned copy is as big as the original. Methods such as LoRA freeze the original weights and train small add-on pieces instead.

2 · Why it exists

Training a big model from nothing for every new job is slow and needs lots of data.

Too little dataMany real tasks have too few examples to train a full-size model from scratch.
Too much computeStarting from a pretrained model needs far less computing power, data and time than starting from random weights.
Wrong final layerThe last part of an image model is tied to the list of classes it was first trained on.
3 · How it works

Follow a cat-or-dog photo model as it is fine-tuned.

How fine-tuning adapts a pretrained model A pretrained model and a small set of labelled cat and dog photos feed step one, where the base layers are frozen and only a new classifier head is trained. The highlighted step two unfreezes some base layers and retrains with a very low learning rate. The result is a fine-tuned cat-or-dog model. INPUTS TRAINING · TWO ROUNDS OUTPUT Pretrained model Labelled photos 1. New head 2. Unfreeze Fine-tuned model weights already learned elsewhere small set of cats and dogs base layers frozen train only the new final layer new head, base weights adjusted all or part of the base trains again with small weight changes very low learning rate photo → cat | dog
The model keeps what it learned before. Small, careful weight changes adapt it to the new task.
  1. 1 · loadLoad a model whose weights were already learned on a large dataset.
  2. 2 · freezeFreeze its layers, so their weights do not change during training.
  3. 3 · headAdd a new final layer for your task and train only that layer.
  4. 4 · unfreezeLet some or all base layers change again, then retrain while keeping the learning rate very low.
  5. 5 · watchKeep weight changes small, because big ones on a small dataset risk overfitting.

Fine-tuning starts from learned weights, not random ones, and changes them gently.

4 · Where it's used
WhoWhat they askWhat it works with
Photo app team“Can a general image model sort cats from dogs”A small set of labelled pet photos
Developer tools team“Can the model get better at writing code”A dataset of coding examples
Support team“Can replies match our style more reliably”Example prompts paired with approved replies
5 · What it solves, and what it doesn't
solves
  • It lets a model reuse what it learned before instead of starting over.
  • It needs far less compute, data and time than training from scratch.
  • It can make a model more reliably produce the style and content you want.
doesn't solve
  • It can overfit quickly when the new dataset is small.
  • Full fine-tuning of a huge model still needs a lot of hardware, and each copy is as large as the original.
  • It does not prove itself; you need tests that compare it with the base model.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsFine-tuning, Hugging Face · read 28 Sept 2026
  2. docsTransfer learning & fine-tuning, Keras · read 28 Sept 2026
  3. docsTransfer learning and fine-tuning, TensorFlow · read 28 Sept 2026
  4. paperUniversal Language Model Fine-tuning for Text Classification, Howard and Ruder · read 28 Sept 2026
  5. paperLoRA: Low-Rank Adaptation of Large Language Models, Hu et al., Microsoft · read 28 Sept 2026
  6. docsSupervised fine-tuning, OpenAI · read 28 Sept 2026