Transformer neural networks
A transformer is a neural network that uses attention to decide which parts of a sequence matter to one another.
A transformer replaces recurrence and convolution with attention as its central sequence mechanism. In an encoder’s self-attention layer, the queries, keys and values all come from the previous layer, so each position consults the others in the same sequence. The result is another representation for each position, now shaped by relevant context elsewhere in the sequence.
Attention does not receive words directly. It receives vectors called queries, keys and values. A query is compared with the keys, and the resulting weights decide how much of each value flows into the output. Multi-head attention repeats that comparison through several learned projections, giving the layer more than one way to connect positions.
A decoder-style transformer hides future positions. It writes its output one piece at a time, feeding each new piece back in. An encoder-style transformer can let each position attend across the whole supplied sequence. Encoder-only and decoder-only architectures are both transformer variants.
BERT shows the encoder-style path, where self-attention uses context on both sides. GPT-3, by contrast, is a 175-billion-parameter autoregressive language model. Its authors gave it each task and a few worked examples as plain text, without updating any of its weights.
Recurrent models, an earlier design, typically work through a sequence one position at a time.
Follow one word through a transformer layer.
- 1 · embedEach token becomes a learned vector, and position information is added so order is not lost.
- 2 · compareEach word makes a query, and the query is scored against the key of each word. That includes the word's own key.
- 3 · attendThe comparison scores become weights used to mix the corresponding value vectors.
- 4 · combineSeveral attention heads perform this work in parallel and their results are joined.
- 5 · refineA feed-forward network then reworks each position on its own. The original model stacks six of these layers, each feeding the next.
The key move is self-attention: a position builds its new representation from weighted information at other positions.
| Who | What they ask | What it works with |
|---|---|---|
| Translation system | “What does this sentence mean in French?” | Every token and its surrounding context |
| Writing assistant | “What should come after this paragraph?” | The prompt and text already generated |
| Document classifier | “Which category fits this report?” | The words and their relationships |
| Vision model | “Which image regions belong together?” | Patches of the image |
- In an encoder, any word can pull in information from any other word in the input.
- Splitting attention into several heads lets one layer look at the words in several learned ways at once, instead of one blurred average.
- With no step-by-step chain, far more of the training work can happen at the same time.
- A mask in the decoder hides later words, so the model can learn to guess the next word without peeking.
- Making a language model bigger does not by itself make it better at doing what you asked.
- Attention on its own has no sense of word order, so position information has to be added in.
- Very long inputs are costly. The original paper suggests letting each word look only at a nearby window when sequences get very long.
- Language models built this way can still write text that is false, harmful or useless to the person asking.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperAttention Is All You Need, Vaswani et al. · read 27 Sept 2026
- docsLLMs: What's a large language model?, Google for Developers · read 27 Sept 2026
- paperTraining language models to follow instructions with human feedback, Ouyang et al. · read 27 Sept 2026
- paperLanguage Models are Few-Shot Learners, Brown et al. · read 27 Sept 2026
- paperBERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, Devlin et al. · read 27 Sept 2026