Transfer learning
Transfer learning reuses representations learned on a source task to improve learning on a different target task.
Transfer learning splits training into two stages. A network first learns from a big source task, and those learned weights then become the starting point for a smaller target task. The R-CNN object detector used this recipe because labelled detection data was scarce. For text, ULMFiT did the same with a language model and used 100 labelled examples to match a from-scratch model trained on 100 times more data.
Not every layer carries over equally well. In image networks, the first layers pick up simple edge-like filters and colour blobs that look alike from one dataset to the next. The top layers, by contrast, tune themselves to the original task. So the old output layer is swapped for a new one, and the team decides whether the copied layers stay frozen or keep training. Freezing is the safer choice with a tiny dataset, because training many parameters on few examples can overfit.
Transfer can also go wrong. Features help less as the target task drifts away from the source. Fine-tuning too hard can erase what pretraining learned, so ULMFiT unfreezes layers gradually to keep that knowledge. And a model pretrained on one large dataset still carries some of that dataset’s bias. Domain-adaptation methods such as Deep Domain Confusion add an extra layer and a loss that push source and target features to look alike.
Target datasets are often too small or costly to support training a capable model from random initialization.
Reuse a pretrained backbone, then adapt it to the target.
- 1 · pretrainLearn general representations from a large source dataset or objective.
- 2 · replaceReplace the source output layer with one shaped for the target task.
- 3 · adaptKeep the copied layers frozen, or let them keep training on the target data.
- 4 · compareCheck the result against a model trained from scratch on the target task.
Transfer is valuable when the source representation is relevant enough; similarity is an empirical question, not a guarantee.
| Who | What they ask | What it works with |
|---|---|---|
| Applied scientist | “Which checkpoint is closest to my data?” | Target validation results by source model |
| Trainer | “Which layers should remain frozen?” | Layer-wise fine-tuning ablations |
| Evaluator | “Does the model still work in the new domain?” | Held-out data from the target domain |
- Transfer learning can reduce target-label requirements.
- Pretrained features can improve generalization after target fine-tuning.
- One source model can initialize many downstream tasks.
- Distant source and target tasks can transfer poorly.
- Splitting a network can break layers that learned to work together.
- Fine-tuning too aggressively can erase what pretraining learned.
- Pretraining on a large dataset reduces dataset bias but does not remove it.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperHow transferable are features in deep neural networks?, Yosinski et al. · read 27 Sept 2026
- paperUniversal Language Model Fine-tuning for Text Classification, Howard and Ruder · read 27 Sept 2026
- paperBERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, Devlin et al. · read 27 Sept 2026
- paperRich feature hierarchies for accurate object detection and semantic segmentation, Girshick et al. · read 27 Sept 2026
- paperDeep Domain Confusion: Maximizing for Domain Invariance, Tzeng et al. · read 27 Sept 2026