Encoder-only modelsConcepts

Encoder-only transformer models

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

An encoder-only model uses bidirectional self-attention to represent a supplied sequence.

1 · What it is

BERT is a bidirectional Transformer encoder. Its self-attention can use context on both sides of each supplied token. BERT uses absolute position embeddings.

Its pre-training masks some input tokens and predicts them. BERT can be fine-tuned with one additional output layer.

RoBERTa is a replication study of BERT pre-training. ELECTRA trains a discriminator to identify replaced tokens. DistilBERT uses knowledge distillation during pre-training.

2 · Why it exists

Understanding a supplied sequence can require context from both sides of a position.

Two-sided contextBERT jointly conditions on left and right context in every layer.
Pre-trainingBERT pre-trains deep bidirectional representations from unlabeled text.
Task transferA pre-trained BERT model can be fine-tuned with one additional output layer.
3 · How it works

Follow one masked token through an encoder-only model.

Every supplied position can shape the representation at the masked position. A task head reads the resulting encoder state.
  1. 1 · embedEach input combines token, segment and position embeddings.
  2. 2 · attendBidirectional self-attention uses left and right context.
  3. 3 · encodeBERT assigns a final hidden vector to each input token.
  4. 4 · readBERT can be fine-tuned with one additional output layer.

ELECTRA pre-trains Transformer text encoders.

4 · Where it's used
WhoWhat they askWhat it works with
Search team“Which passages are relevant to this query?”Contextual representations of the query and document
Support team“Which intent label fits this message?”The representation of the supplied message
Extraction system“Which tokens name a person or place?”One contextual representation per token
5 · What it solves, and what it doesn't
solves
  • Bidirectional conditioning lets a position use left and right context.
  • One added output layer can adapt BERT to several language understanding tasks.
  • ELECTRA pre-trains a discriminator to label every token as original or replaced.
doesn't solve
  • BERT uses masked language modeling rather than left-to-right language modeling.
  • Masked language modeling predicts only the selected masked positions during pre-training.