Concepts

Pretraining

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

In one transfer-learning setup, a model is pretrained on a data-rich task before fine-tuning on a downstream task.

1 · What it is

GPT-1 combines generative pretraining on unlabelled text with task-specific fine-tuning.

Pretraining objectives differ. BERT jointly conditions on left and right context in all layers. BART uses a denoising objective that corrupts text and reconstructs the original. ELECTRA trains a discriminator to decide whether each input token was replaced.

T5 compares pretraining objectives, architectures, unlabelled datasets and transfer approaches.

2 · Why it exists

Task labels can be scarce even when unlabelled data is abundant.

Few labelsGPT-1 notes that labelled data for specific language tasks is scarce.
Shared foundationT5 studies transfer learning by pretraining on a data-rich task before fine-tuning on a downstream task.
Different objectivesBART learns to reconstruct text after a noising function corrupts it.
3 · How it works

Follow model state from a broad objective to a specific task.

A model is pretrained on a data-rich task before fine-tuning on a downstream task.
  1. 1 · collectBegin with a data-rich task.
  2. 2 · pretrainCompare candidate pretraining objectives.
  3. 3 · transferCarry the pretrained model into a downstream task.
  4. 4 · adaptFine-tune the model with the task-specific training setup.

T5 compares several pretraining objectives.

4 · Where it's used
WhoWhat they askWhat it works with
Language-model team“Which broad objective should run before task adaptation?”The pretraining objective and corpus
Applied NLP team“Which pretrained checkpoint should initialize this classifier?”A compatible model and tokenizer
Research engineer“Does masking, denoising or next-token prediction fit this architecture?”The objective used during pretraining
Evaluation team“What changes after task-specific fine-tuning?”Results before and after adaptation
5 · What it solves, and what it doesn't
solves
  • Generative pretraining can use a diverse corpus of unlabelled text before task-specific fine-tuning.
  • BERT pretrains bidirectional representations from unlabelled text.
  • BART pretrains by corrupting text and reconstructing the original.
  • ELECTRA pretrains with replaced-token detection.
doesn't solve
  • Hyperparameter choices can significantly affect pretraining results.
  • Downstream tasks still use fine-tuning in the transfer-learning setup described by T5.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperImproving Language Understanding by Generative Pre-Training, OpenAI · read 28 Sept 2026
  2. paperBERT Pre-training of Deep Bidirectional Transformers for Language Understanding, Devlin and colleagues · read 28 Sept 2026
  3. paperBART Denoising Sequence-to-Sequence Pre-training, Lewis and colleagues · read 28 Sept 2026
  4. paperExploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, Raffel and colleagues, JMLR · read 28 Sept 2026
  5. paperELECTRA Pre-training Text Encoders as Discriminators Rather Than Generators, Clark and colleagues · read 28 Sept 2026
  6. paperRoBERTa A Robustly Optimized BERT Pretraining Approach, Liu and colleagues · read 28 Sept 2026