Concepts

Transfer learning

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

Transfer learning reuses representations learned on a source task to improve learning on a different target task.

1 · What it is

Transfer learning splits training into two stages. A network first learns from a big source task, and those learned weights then become the starting point for a smaller target task. The R-CNN object detector used this recipe because labelled detection data was scarce. For text, ULMFiT did the same with a language model and used 100 labelled examples to match a from-scratch model trained on 100 times more data.

Not every layer carries over equally well. In image networks, the first layers pick up simple edge-like filters and colour blobs that look alike from one dataset to the next. The top layers, by contrast, tune themselves to the original task. So the old output layer is swapped for a new one, and the team decides whether the copied layers stay frozen or keep training. Freezing is the safer choice with a tiny dataset, because training many parameters on few examples can overfit.

Transfer can also go wrong. Features help less as the target task drifts away from the source. Fine-tuning too hard can erase what pretraining learned, so ULMFiT unfreezes layers gradually to keep that knowledge. And a model pretrained on one large dataset still carries some of that dataset’s bias. Domain-adaptation methods such as Deep Domain Confusion add an extra layer and a loss that push source and target features to look alike.

2 · Why it exists

Target datasets are often too small or costly to support training a capable model from random initialization.

Sparse labelsA target task may have only a hundred or so labelled examples.
Repeated workThe first layers of image networks learn similar simple patterns whatever the dataset, so starting from zero rebuilds them.
Domain gapSource-specific features can become less useful as tasks diverge.
3 · How it works

Reuse a pretrained backbone, then adapt it to the target.

Transfer learning reuses general layers and replaces the task headA large source dataset trains general layers and a source-specific head. The general layers are copied into the target task, where a small labelled dataset feeds them. The highlighted step adds a new target head and fine-tunes the copied layers, producing a target model checked on unseen data.SOURCE TASKLarge datasetlots of labelsGeneral layersearly featuresSource headsource classescopy weightsTARGET TASKSmall datasetfew labelsCopied layersfreeze or tunereused featuresKEY STEPAdapt to targetnew target head +tune copied layersTarget modelcheck on unseen data
Transfer keeps reusable representations and relearns the task-specific output.
  1. 1 · pretrainLearn general representations from a large source dataset or objective.
  2. 2 · replaceReplace the source output layer with one shaped for the target task.
  3. 3 · adaptKeep the copied layers frozen, or let them keep training on the target data.
  4. 4 · compareCheck the result against a model trained from scratch on the target task.

Transfer is valuable when the source representation is relevant enough; similarity is an empirical question, not a guarantee.

4 · Where it's used
WhoWhat they askWhat it works with
Applied scientist“Which checkpoint is closest to my data?”Target validation results by source model
Trainer“Which layers should remain frozen?”Layer-wise fine-tuning ablations
Evaluator“Does the model still work in the new domain?”Held-out data from the target domain
5 · What it solves, and what it doesn't
solves
  • Transfer learning can reduce target-label requirements.
  • Pretrained features can improve generalization after target fine-tuning.
  • One source model can initialize many downstream tasks.
doesn't solve
  • Distant source and target tasks can transfer poorly.
  • Splitting a network can break layers that learned to work together.
  • Fine-tuning too aggressively can erase what pretraining learned.
  • Pretraining on a large dataset reduces dataset bias but does not remove it.