Inference optimization
Inference optimization is a set of methods, like caching, quantization and speculative decoding, that let a trained language model answer with less work or less memory.
A language model writes one token (a word or piece of a word) at a time, which can make generation slow. Large models also need a lot of GPU memory just to run. Inference optimization is the toolbox for doing the same job with less time and memory.
Some tools cut memory. Quantization stores the weights (the numbers the model learned in training) in smaller, less precise numbers. It is a bit like saving a photo at lower quality so it takes less space. PagedAttention, used by the vLLM server, manages the model’s working cache with an idea borrowed from how operating systems handle memory.
Other tools cut repeated work. A key-value cache saves results for earlier tokens so they are not worked out again. FlashAttention gives the same exact attention answer but moves less data between two kinds of memory on a GPU (the chip that runs the model): the GPU’s main high-bandwidth memory and its on-chip memory. Speculative decoding lets a small model guess a few tokens ahead. It needs no retraining and no change to the model’s design.
Running a large language model is slow and memory-hungry.
Follow one reply through an optimized model.
- 1 · shrinkQuantization stores the weights in lower-precision numbers so the model needs less memory.
- 2 · cacheA key-value cache stores attention results from earlier tokens and reuses them.
- 3 · draftA small, fast model guesses the next few tokens.
- 4 · verifyThe large model checks all the guesses at once and keeps the ones that fit its own choices, so the output follows the same odds as before.
The biggest saving is not redoing work the model has already done.
| Who | What they ask | What it works with |
|---|---|---|
| Serving team | “How many users can one GPU handle at once?” | KV cache memory and batch size |
| App developer | “Can this model fit on a smaller GPU?” | 8-bit quantization of the weights |
| Chat product team | “Can replies stream faster without changing them?” | Speculative decoding with a draft model |
- A KV cache avoids recomputing attention for earlier tokens.
- 8-bit weights can halve inference memory in the LLM.int8() method.
- Speculative decoding can speed up generation without changing outputs.
- vLLM achieves near-zero waste in KV cache memory.
- Quantization tries to keep accuracy but stores weights less precisely.
- A KV cache trades compute for memory that grows with the sequence.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsHow caching works, Hugging Face · read 28 Sept 2026
- docsQuantization overview, Hugging Face · read 28 Sept 2026
- paperLLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale, arXiv · read 28 Sept 2026
- paperFast Inference from Transformers via Speculative Decoding, arXiv · read 28 Sept 2026
- paperEfficient Memory Management for Large Language Model Serving with PagedAttention, arXiv · read 28 Sept 2026
- paperFlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness, arXiv · read 28 Sept 2026