Encoder-only modelsConcepts
Encoder-only transformer models
1 · In one line
An encoder-only model uses bidirectional self-attention to represent a supplied sequence.
1 · What it is
BERT is a bidirectional Transformer encoder. Its self-attention can use context on both sides of each supplied token. BERT uses absolute position embeddings.
Its pre-training masks some input tokens and predicts them. BERT can be fine-tuned with one additional output layer.
RoBERTa is a replication study of BERT pre-training. ELECTRA trains a discriminator to identify replaced tokens. DistilBERT uses knowledge distillation during pre-training.
Understanding a supplied sequence can require context from both sides of a position.
Two-sided contextBERT jointly conditions on left and right context in every layer.
Pre-trainingBERT pre-trains deep bidirectional representations from unlabeled text.
Task transferA pre-trained BERT model can be fine-tuned with one additional output layer.
Follow one masked token through an encoder-only model.
- 1 · embedEach input combines token, segment and position embeddings.
- 2 · attendBidirectional self-attention uses left and right context.
- 3 · encodeBERT assigns a final hidden vector to each input token.
- 4 · readBERT can be fine-tuned with one additional output layer.
ELECTRA pre-trains Transformer text encoders.
| Who | What they ask | What it works with |
|---|---|---|
| Search team | “Which passages are relevant to this query?” | Contextual representations of the query and document |
| Support team | “Which intent label fits this message?” | The representation of the supplied message |
| Extraction system | “Which tokens name a person or place?” | One contextual representation per token |
solves
- Bidirectional conditioning lets a position use left and right context.
- One added output layer can adapt BERT to several language understanding tasks.
- ELECTRA pre-trains a discriminator to label every token as original or replaced.
doesn't solve
- BERT uses masked language modeling rather than left-to-right language modeling.
- Masked language modeling predicts only the selected masked positions during pre-training.
6 · Go deeper
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperBERT Pre-training of Deep Bidirectional Transformers for Language Understanding, Devlin et al. · read 28 Sept 2026
- paperRoBERTa A Robustly Optimized BERT Pretraining Approach, Liu et al. · read 28 Sept 2026
- paperELECTRA Pre-training Text Encoders as Discriminators Rather Than Generators, Clark et al. · read 28 Sept 2026
- paperDistilBERT a distilled version of BERT smaller faster cheaper and lighter, Sanh et al. · read 28 Sept 2026
- docsBERT, Hugging Face · read 28 Sept 2026