LLMConcepts

Large language model

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

An LLM predicts a token or sequence of tokens, sometimes many paragraphs long.

1 · What it is

An LLM is a neural network. It predicts a token or a sequence of tokens. Training an LLM from scratch uses an enormous collection of text. A token may be a whole word, part of a word or one character.

BERT-style masked training predicts selected hidden tokens using context on both sides. Autoregressive training hides future positions so each prediction depends only on earlier output.

Tasks can be specified through text interaction rather than task-specific gradient updates. During masked training, loss guides backpropagation as it updates parameter values. Through that practice, a transformer-based model can learn patterns and higher-order structures in the data. LLMs contain many parameters and can gather substantial context. Large language models can still generate untruthful outputs.

Instruction tuning is an optional further training step that can improve instruction following. Instruction-tuned systems may use demonstrations and human rankings of candidate outputs. InstructGPT still makes simple mistakes.

2 · Why it exists

Text interaction can specify many tasks to one language model.

Many formsA token can be a word, part of a word or even one character.
Many tasksThe same model can be prompted for translation, question answering, completion and other language tasks.
Intent gapMaking a language model larger does not by itself make the model follow a user's intent better.
3 · How it works

In masked pretraining, selected tokens are hidden and the model tries to predict them.

  1. 1 · tokeniseTraining text is divided into tokens, which may be words, word pieces or individual characters.
  2. 2 · predictIn masked training, the model repeatedly tries to predict tokens hidden from its input.
  3. 3 · adjustLoss guides backpropagation as it updates the model's parameter values.
  4. 4 · alignDemonstrations and human rankings can be used to fine-tune a pretrained model to follow instructions.
  5. 5 · generateAn LLM can predict token sequences many paragraphs long.

An LLM generates probabilities for possible responses.

4 · Where it's used
WhoWhat they askWhat it works with
Analyst“Summarise the differences between these reports.”Prompt, supplied reports and learned language patterns
Developer“Write a test for this function.”Instructions, code and surrounding context
Support team“Draft a clear reply to this customer.”Customer message, policy context and response instructions
Student“Explain this equation in simpler words.”The equation, prompt and model context
5 · What it solves, and what it doesn't
solves
  • One pretrained language model can attempt many tasks through text instructions and examples.
  • Large language models can generate long sequences rather than predicting only one token.
  • GPT-3 was applied to all evaluated tasks without gradient updates or fine-tuning at use time.
  • Human preferences can be used as a reward signal to fine-tune a model.
doesn't solve
  • Large language models can generate untruthful outputs.
  • A larger model does not automatically follow a user's intent.
  • Language models can produce text that is untruthful, toxic or unhelpful.
  • Few-shot performance still varies by task, even for a very large model.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsLLMs: What's a large language model?, Google for Developers · read 28 Sept 2026
  2. paperLanguage Models are Few-Shot Learners, Brown et al. · read 28 Sept 2026
  3. paperTraining language models to follow instructions with human feedback, Ouyang et al. · read 28 Sept 2026
  4. paperAttention Is All You Need, Vaswani et al. · read 28 Sept 2026
  5. paperBERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, Devlin et al. · read 28 Sept 2026