Large language model
An LLM predicts a token or sequence of tokens, sometimes many paragraphs long.
An LLM is a neural network. It predicts a token or a sequence of tokens. Training an LLM from scratch uses an enormous collection of text. A token may be a whole word, part of a word or one character.
BERT-style masked training predicts selected hidden tokens using context on both sides. Autoregressive training hides future positions so each prediction depends only on earlier output.
Tasks can be specified through text interaction rather than task-specific gradient updates. During masked training, loss guides backpropagation as it updates parameter values. Through that practice, a transformer-based model can learn patterns and higher-order structures in the data. LLMs contain many parameters and can gather substantial context. Large language models can still generate untruthful outputs.
Instruction tuning is an optional further training step that can improve instruction following. Instruction-tuned systems may use demonstrations and human rankings of candidate outputs. InstructGPT still makes simple mistakes.
Text interaction can specify many tasks to one language model.
In masked pretraining, selected tokens are hidden and the model tries to predict them.
- 1 · tokeniseTraining text is divided into tokens, which may be words, word pieces or individual characters.
- 2 · predictIn masked training, the model repeatedly tries to predict tokens hidden from its input.
- 3 · adjustLoss guides backpropagation as it updates the model's parameter values.
- 4 · alignDemonstrations and human rankings can be used to fine-tune a pretrained model to follow instructions.
- 5 · generateAn LLM can predict token sequences many paragraphs long.
An LLM generates probabilities for possible responses.
| Who | What they ask | What it works with |
|---|---|---|
| Analyst | “Summarise the differences between these reports.” | Prompt, supplied reports and learned language patterns |
| Developer | “Write a test for this function.” | Instructions, code and surrounding context |
| Support team | “Draft a clear reply to this customer.” | Customer message, policy context and response instructions |
| Student | “Explain this equation in simpler words.” | The equation, prompt and model context |
- One pretrained language model can attempt many tasks through text instructions and examples.
- Large language models can generate long sequences rather than predicting only one token.
- GPT-3 was applied to all evaluated tasks without gradient updates or fine-tuning at use time.
- Human preferences can be used as a reward signal to fine-tune a model.
- Large language models can generate untruthful outputs.
- A larger model does not automatically follow a user's intent.
- Language models can produce text that is untruthful, toxic or unhelpful.
- Few-shot performance still varies by task, even for a very large model.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsLLMs: What's a large language model?, Google for Developers · read 28 Sept 2026
- paperLanguage Models are Few-Shot Learners, Brown et al. · read 28 Sept 2026
- paperTraining language models to follow instructions with human feedback, Ouyang et al. · read 28 Sept 2026
- paperAttention Is All You Need, Vaswani et al. · read 28 Sept 2026
- paperBERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, Devlin et al. · read 28 Sept 2026