Concepts

Inference optimization

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

Inference optimization is a set of methods, like caching, quantization and speculative decoding, that let a trained language model answer with less work or less memory.

1 · What it is

A language model writes one token (a word or piece of a word) at a time, which can make generation slow. Large models also need a lot of GPU memory just to run. Inference optimization is the toolbox for doing the same job with less time and memory.

Some tools cut memory. Quantization stores the weights (the numbers the model learned in training) in smaller, less precise numbers. It is a bit like saving a photo at lower quality so it takes less space. PagedAttention, used by the vLLM server, manages the model’s working cache with an idea borrowed from how operating systems handle memory.

Other tools cut repeated work. A key-value cache saves results for earlier tokens so they are not worked out again. FlashAttention gives the same exact attention answer but moves less data between two kinds of memory on a GPU (the chip that runs the model): the GPU’s main high-bandwidth memory and its on-chip memory. Speculative decoding lets a small model guess a few tokens ahead. It needs no retraining and no change to the model’s design.

2 · Why it exists

Running a large language model is slow and memory-hungry.

One word at a timeDecoding K tokens takes K runs of the model, one after another.
Big models, big memoryLarge language models need a lot of GPU memory just to run.
Wasted cache spaceServing systems can waste the memory each request keeps for its past work, which limits how many requests share the GPU.
3 · How it works

Follow one reply through an optimized model.

Each trick removes repeated work or memory; caching keeps outputs the same and speculative decoding keeps them statistically the same, while quantization saves memory by storing weights less precisely.
  1. 1 · shrinkQuantization stores the weights in lower-precision numbers so the model needs less memory.
  2. 2 · cacheA key-value cache stores attention results from earlier tokens and reuses them.
  3. 3 · draftA small, fast model guesses the next few tokens.
  4. 4 · verifyThe large model checks all the guesses at once and keeps the ones that fit its own choices, so the output follows the same odds as before.

The biggest saving is not redoing work the model has already done.

4 · Where it's used
WhoWhat they askWhat it works with
Serving team“How many users can one GPU handle at once?”KV cache memory and batch size
App developer“Can this model fit on a smaller GPU?”8-bit quantization of the weights
Chat product team“Can replies stream faster without changing them?”Speculative decoding with a draft model
5 · What it solves, and what it doesn't
solves
  • A KV cache avoids recomputing attention for earlier tokens.
  • 8-bit weights can halve inference memory in the LLM.int8() method.
  • Speculative decoding can speed up generation without changing outputs.
  • vLLM achieves near-zero waste in KV cache memory.
doesn't solve
  • Quantization tries to keep accuracy but stores weights less precisely.
  • A KV cache trades compute for memory that grows with the sequence.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsHow caching works, Hugging Face · read 28 Sept 2026
  2. docsQuantization overview, Hugging Face · read 28 Sept 2026
  3. paperLLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale, arXiv · read 28 Sept 2026
  4. paperFast Inference from Transformers via Speculative Decoding, arXiv · read 28 Sept 2026
  5. paperEfficient Memory Management for Large Language Model Serving with PagedAttention, arXiv · read 28 Sept 2026
  6. paperFlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness, arXiv · read 28 Sept 2026