Concepts

Inference

5 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Inference is using a trained AI model on new input to get an answer, such as a label, a number or a written reply, without changing what the model learned.

1 · What it is

Inference is the stage where a trained model gets used. You give it an input it has not seen before, and it produces an output: a label for a photo, a number, or a written reply. Training has already set the model’s weights, the learned numbers inside it. During inference those weights are only read, never changed, and they stay loaded in memory while the model serves requests. (In statistics, “inference” means something different, so check which sense a text uses.)

For a simple classifier, one forward pass, a single trip of the input through the model, gives the answer. A language model works in two phases. In prefill, it reads the whole prompt at once, split into tokens of about four English characters each, and stores intermediate results called keys and values. That same pass also yields the first token of the reply. In decode, it adds one token per pass until it hits a limit or an end token. A key-value (KV) cache keeps the stored results so each step does not redo earlier work.

That loop is why speed and cost matter. Generating tokens is almost always the slowest step, so halving the length of a reply can roughly halve the wait, while halving the prompt may save only 1 to 5 per cent. Apps track time to first token and stream text as it is made. Serving large models also needs many GPUs, which makes each request costly.

So engineers choose how and where to run it. Static, or batch, inference computes answers ahead of time, like a weather model that refreshes forecasts every four hours. Dynamic, or online, inference runs on each request, handling any input but needing to be fast. Batching requests raises throughput at some cost to waiting time; Anthropic’s batch API halves the price for jobs that can wait. Quantisation, storing weights with fewer bits, can make a model four times smaller so it runs on a phone, offline, keeping data on the device.

2 · Why it exists

A trained model is only useful once it answers real requests, and that has costs.

People are waitingLatency, the time a model takes to process input and respond, is a real concern when answers are made on demand.
Every answer uses hardwareRunning large language models is expensive and needs many accelerator chips such as GPUs.
Replies come out in piecesA language model writes one token at a time, and generating those tokens is usually the slowest part.
3 · How it works

Follow one chat request through a language model.

The prompt is read once; the reply is built one token per pass. Output length drives most of the wait.
  1. 1 · tokeniseThe prompt is split into tokens, small chunks of text of roughly four English characters each.
  2. 2 · prefillThe model reads all the prompt tokens in one parallel pass, stores intermediate results called keys and values, and works out the first token of the reply.
  3. 3 · decodeIt then produces one new token per pass, reusing the stored keys and values, until a stop condition is met.
  4. 4 · streamTokens are turned back into text and can be streamed, so the reader sees the reply before it is finished.

The weights do not change during inference. Training set them; every request only reads them.

4 · Where it's used
WhoWhat they askWhat it works with
Chat app team“How long does a user wait before the first word of a reply appears?”Time to first token and tokens per second for each request
Weather service“Should we compute forecasts every few hours or run the model on every request?”Static versus dynamic inference for the forecast model
Phone app developer“Can this feature work with no network and keep photos on the phone?”A quantised model running on the device
Research data team“Can we label a million documents overnight for less money?”Batch jobs sent to a model API
5 · What it solves, and what it doesn't
solves
  • Turns a trained model into answers for new inputs it has never seen.
  • Online inference can answer any new input as it arrives, including rare ones.
  • Static inference lets predictions be checked before anyone uses them.
  • Caching, batching and quantisation make each answer cheaper or faster.
doesn't solve
  • It does not teach the model anything; the weights stay exactly as training left them.
  • Static inference cannot answer an input that was not computed in advance.
  • Adding more servers does not necessarily fix slow responses.
  • Quantisation makes a model smaller but can cost a little accuracy.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
  2. docsProduction ML systems: Static versus dynamic inference (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  3. officialMastering LLM Techniques: Inference Optimization, NVIDIA Technical Blog · read 27 Sept 2026
  4. paperEfficient Memory Management for Large Language Model Serving with PagedAttention, arXiv (UC Berkeley, Kwon et al.; SOSP 2023) · read 27 Sept 2026
  5. docsHow caching works (Transformers documentation), Hugging Face · read 27 Sept 2026
  6. docsBatchers (Triton Inference Server user guide), NVIDIA · read 27 Sept 2026
  7. docsLatency optimization, OpenAI · read 27 Sept 2026
  8. docsReducing latency, Anthropic · read 27 Sept 2026
  9. docsBatch processing, Anthropic · read 27 Sept 2026
  10. docsGetting started with LiteRT, Google AI Edge · read 27 Sept 2026
  11. docsPost-training quantization (LiteRT), Google AI Edge · read 27 Sept 2026
  12. officialML Kit, Google for Developers · read 27 Sept 2026