Inference
Inference is using a trained AI model on new input to get an answer, such as a label, a number or a written reply, without changing what the model learned.
Inference is the stage where a trained model gets used. You give it an input it has not seen before, and it produces an output: a label for a photo, a number, or a written reply. Training has already set the model’s weights, the learned numbers inside it. During inference those weights are only read, never changed, and they stay loaded in memory while the model serves requests. (In statistics, “inference” means something different, so check which sense a text uses.)
For a simple classifier, one forward pass, a single trip of the input through the model, gives the answer. A language model works in two phases. In prefill, it reads the whole prompt at once, split into tokens of about four English characters each, and stores intermediate results called keys and values. That same pass also yields the first token of the reply. In decode, it adds one token per pass until it hits a limit or an end token. A key-value (KV) cache keeps the stored results so each step does not redo earlier work.
That loop is why speed and cost matter. Generating tokens is almost always the slowest step, so halving the length of a reply can roughly halve the wait, while halving the prompt may save only 1 to 5 per cent. Apps track time to first token and stream text as it is made. Serving large models also needs many GPUs, which makes each request costly.
So engineers choose how and where to run it. Static, or batch, inference computes answers ahead of time, like a weather model that refreshes forecasts every four hours. Dynamic, or online, inference runs on each request, handling any input but needing to be fast. Batching requests raises throughput at some cost to waiting time; Anthropic’s batch API halves the price for jobs that can wait. Quantisation, storing weights with fewer bits, can make a model four times smaller so it runs on a phone, offline, keeping data on the device.
A trained model is only useful once it answers real requests, and that has costs.
Follow one chat request through a language model.
- 1 · tokeniseThe prompt is split into tokens, small chunks of text of roughly four English characters each.
- 2 · prefillThe model reads all the prompt tokens in one parallel pass, stores intermediate results called keys and values, and works out the first token of the reply.
- 3 · decodeIt then produces one new token per pass, reusing the stored keys and values, until a stop condition is met.
- 4 · streamTokens are turned back into text and can be streamed, so the reader sees the reply before it is finished.
The weights do not change during inference. Training set them; every request only reads them.
| Who | What they ask | What it works with |
|---|---|---|
| Chat app team | “How long does a user wait before the first word of a reply appears?” | Time to first token and tokens per second for each request |
| Weather service | “Should we compute forecasts every few hours or run the model on every request?” | Static versus dynamic inference for the forecast model |
| Phone app developer | “Can this feature work with no network and keep photos on the phone?” | A quantised model running on the device |
| Research data team | “Can we label a million documents overnight for less money?” | Batch jobs sent to a model API |
- Turns a trained model into answers for new inputs it has never seen.
- Online inference can answer any new input as it arrives, including rare ones.
- Static inference lets predictions be checked before anyone uses them.
- Caching, batching and quantisation make each answer cheaper or faster.
- It does not teach the model anything; the weights stay exactly as training left them.
- Static inference cannot answer an input that was not computed in advance.
- Adding more servers does not necessarily fix slow responses.
- Quantisation makes a model smaller but can cost a little accuracy.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
- docsProduction ML systems: Static versus dynamic inference (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- officialMastering LLM Techniques: Inference Optimization, NVIDIA Technical Blog · read 27 Sept 2026
- paperEfficient Memory Management for Large Language Model Serving with PagedAttention, arXiv (UC Berkeley, Kwon et al.; SOSP 2023) · read 27 Sept 2026
- docsHow caching works (Transformers documentation), Hugging Face · read 27 Sept 2026
- docsBatchers (Triton Inference Server user guide), NVIDIA · read 27 Sept 2026
- docsLatency optimization, OpenAI · read 27 Sept 2026
- docsReducing latency, Anthropic · read 27 Sept 2026
- docsBatch processing, Anthropic · read 27 Sept 2026
- docsGetting started with LiteRT, Google AI Edge · read 27 Sept 2026
- docsPost-training quantization (LiteRT), Google AI Edge · read 27 Sept 2026
- officialML Kit, Google for Developers · read 27 Sept 2026