TransformersConcepts

Transformer neural networks

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

A transformer is a neural network that uses attention to decide which parts of a sequence matter to one another.

1 · What it is

A transformer replaces recurrence and convolution with attention as its central sequence mechanism. In an encoder’s self-attention layer, the queries, keys and values all come from the previous layer, so each position consults the others in the same sequence. The result is another representation for each position, now shaped by relevant context elsewhere in the sequence.

Attention does not receive words directly. It receives vectors called queries, keys and values. A query is compared with the keys, and the resulting weights decide how much of each value flows into the output. Multi-head attention repeats that comparison through several learned projections, giving the layer more than one way to connect positions.

A decoder-style transformer hides future positions. It writes its output one piece at a time, feeding each new piece back in. An encoder-style transformer can let each position attend across the whole supplied sequence. Encoder-only and decoder-only architectures are both transformer variants.

BERT shows the encoder-style path, where self-attention uses context on both sides. GPT-3, by contrast, is a 175-billion-parameter autoregressive language model. Its authors gave it each task and a few worked examples as plain text, without updating any of its weights.

2 · Why it exists

Recurrent models, an earlier design, typically work through a sequence one position at a time.

Long relayA recurrent model works out its memory for each word from its memory of the word before.
Less parallel workBecause each step waits for the last, the words of one training example cannot all be handled at the same moment.
Distant contextAttention links two words directly, however many words sit between them.
3 · How it works

Follow one word through a transformer layer.

  1. 1 · embedEach token becomes a learned vector, and position information is added so order is not lost.
  2. 2 · compareEach word makes a query, and the query is scored against the key of each word. That includes the word's own key.
  3. 3 · attendThe comparison scores become weights used to mix the corresponding value vectors.
  4. 4 · combineSeveral attention heads perform this work in parallel and their results are joined.
  5. 5 · refineA feed-forward network then reworks each position on its own. The original model stacks six of these layers, each feeding the next.

The key move is self-attention: a position builds its new representation from weighted information at other positions.

4 · Where it's used
WhoWhat they askWhat it works with
Translation system“What does this sentence mean in French?”Every token and its surrounding context
Writing assistant“What should come after this paragraph?”The prompt and text already generated
Document classifier“Which category fits this report?”The words and their relationships
Vision model“Which image regions belong together?”Patches of the image
5 · What it solves, and what it doesn't
solves
  • In an encoder, any word can pull in information from any other word in the input.
  • Splitting attention into several heads lets one layer look at the words in several learned ways at once, instead of one blurred average.
  • With no step-by-step chain, far more of the training work can happen at the same time.
  • A mask in the decoder hides later words, so the model can learn to guess the next word without peeking.
doesn't solve
  • Making a language model bigger does not by itself make it better at doing what you asked.
  • Attention on its own has no sense of word order, so position information has to be added in.
  • Very long inputs are costly. The original paper suggests letting each word look only at a nearby window when sequences get very long.
  • Language models built this way can still write text that is false, harmful or useless to the person asking.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperAttention Is All You Need, Vaswani et al. · read 27 Sept 2026
  2. docsLLMs: What's a large language model?, Google for Developers · read 27 Sept 2026
  3. paperTraining language models to follow instructions with human feedback, Ouyang et al. · read 27 Sept 2026
  4. paperLanguage Models are Few-Shot Learners, Brown et al. · read 27 Sept 2026
  5. paperBERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, Devlin et al. · read 27 Sept 2026